EDBT 2026 Demo / reviewers in the wild / expert
Wenbo Zhou 0004
dblp:124/2075-4
· DBLP profile ↗
52ranked-venue papers
3as first author
45since 2021 · last 2026
0000-0002-4703-4641ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 25 since 2021Artificial intelligence and machine learning · 26 · 24 since 2021Security and privacy · 9 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor DisentanglementabstractGenerating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the coarse granularity of existing control mechanisms. To overcome these challenges, we have proposed a novel framework called MF-Speech, which consists of two core components: MF-SpeechEncoder and MF-SpeechGenerator. MF-SpeechEncoder acts as a factor purifier, adopting a multi-objective optimization strategy to decompose the original speech signal into highly pure and independent representations of content, timbre, and emotion. Subsequently, MF-SpeechGenerator functions as a conductor, achieving precise, composable and fine-grained control over these factors through dynamic fusion and Hierarchical Style Adaptive Normalization (HSAN). Experiments demonstrate that in the highly challenging multi-factor compositional speech generation task, MF-Speech significantly outperforms current state-of-the-art methods, achieving a lower word error rate (WER=4.67%), superior style control (SECS=0.5685, Corr=0.68), and the highest subjective evaluation scores (nMOS=3.96, sMOS_t=3.86, sMOS_e=3.78). Furthermore, the learned discrete factors exhibit strong transferability, demonstrating their significant potential as a general-purpose speech representation. Youqing Fang, Pingyu Wu, Guoyang Ye, Wenbo Zhou 0004 |
AAAI | 5 |
| 2026 | EARG-Net: Edge-Aware Reconstruction-Guided Network for Image Manipulation Detection and LocalizationabstractRecent advances in image editing tools, particularly those used in content-aware retouching and object-level manipulation, have raised significant concerns regarding the authenticity of digital images. While many Image Manipulation Detection and Localization (IMDL) methods have been proposed, they often struggle with subtle forgeries, intricate boundary artifacts, and manipulations generated by unseen editing techniques. In this work, we propose a novel edge-aware framework that leverages the strong natural image priors of pre-trained inpainting models to harmonize manipulated regions. By guiding the inpainting process with generated edge-aware masks, our method reconstructs tampered areas using surrounding context, yielding perceptually coherent results. The pixel-wise residual between the original and reconstructed images reveals manipulation-sensitive inconsistencies—particularly around editing boundaries—thereby enabling accurate and generalizable detection and localization. Extensive experiments across multiple benchmarks demonstrate that our approach achieves state-of-the-art performance, especially in challenging scenarios involving realistic and finely retouched image forgeries. Yanpu Yu, Zhaoxin Shi, Tianyi Wei, Wenbo Zhou 0004, Nenghai Yu |
AAAI | 5 |
| 2026 | Trait Activation in Silicon: A Situation-Aware Framework for Psychologically Grounded Role-PlayingabstractRole-playing language models (RPLMs) have made significant strides in mimicking static character identities.However, their personality simulations remain superficial, lacking a profound understanding of complex human psychological mechanisms.We identify a critical bottleneck termed "Personality Inertia"-a behavioral rigidity where RLHF-induced alignment bias traps models in a sanitized, "helpful assistant" persona.This inertia prevents models from adapting to diverse social contexts or expressing essential but negative traits under pressure.To bridge this gap, we propose PD-LLM, a situation-aware framework grounded in Trait Activation Theory.PD-LLM introduces Bipolar Latent Decomposition, which decouples personality traits into bidirectional LoRA adapters.These adapters are dynamically modulated by a situation-aware module based on the DIAMONDS taxonomy, allowing for precise behavioral regulation.Empirical results show that while baseline methods fail to synchronize multidimensional traits under pressure, PD-LLM achieves superior performance in both static fidelity and dynamic adaptability.By advancing from prompt engineering to intrinsic parameter control, PD-LLM effectively overcomes personality rigidity, facilitating the creation of vivid and psychologically consistent agents. Zuolong Li, Pingyu Wu, Xianwen Huang, Tianyi Wei, Wenbo Zhou 0004 |
ACL (1) | 5 |
| 2026 | SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
Zonghao Ying, Aishan Liu, Siyuan Liang 0004, Lei Huang 0015, Jinyang Guo 0002, Wenbo Zhou 0004, Xianglong Liu 0001, Dacheng Tao |
Int. J. Comput. Vis. | 6 |
| 2026 | Unifying Multi-Modal Hair Editing via Proxy Feature BlendingabstractHair editing is a long-standing problem in computer vision that demands both fine-grained local control and intuitive user interactions across diverse modalities. Despite the remarkable progress of GANs and diffusion models, existing methods still lack a unified framework that simultaneously supports arbitrary interaction modes (e.g., text, sketch, mask, and reference image) while ensuring precise editing and faithful preservation of irrelevant attributes. In this work, we introduce a novel paradigm that reformulates hair editing as proxy-based hair transfer. Specifically, we leverage the dense and semantically disentangled latent space of StyleGAN for precise manipulation and exploit its feature space for disentangled attribute preservation, thereby decoupling the objectives of editing and preservation. Our framework unifies different modalities by converting editing conditions into distinct transfer proxies, whose features are seamlessly blended to achieve global or local edits. Beyond 2D, we extend our paradigm to 3D-aware settings by incorporating EG3D and PanoHead, where we propose a multi-view boosted hair feature localization strategy together with 3D-tailored proxy generation methods that exploit the inherent properties of 3D-aware generative models. Extensive experiments demonstrate that our method consistently outperforms prior approaches in editing effects, attribute preservation, visual naturalness, and multi-view consistency, while offering unprecedented support for multimodal and mixed-modal interactions. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Can Wang 0007, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Improving adversarial transferability and imperceptibility with loss landscape and diffusion model
Wenbo Zhou 0004, Ee-Chien Chang, Siew-Kei Lam |
Pattern Recognit. | 4 |
| 2026 | UniForensics: Face Forgery Detection via General Facial RepresentationabstractThe rise of deepfakes has significantly heightened concerns for privacy and the authenticity of digital media, bringing widespread attention to face forgery detection. Previous deepfake detection methods mostly depend on low-level textural features vulnerable to perturbations and fall short of detecting unseen forgery methods. In contrast, high-level semantic features are less susceptible to perturbations and not limited to forgery-specific artifacts, thus having stronger generalization. Motivated by this, we propose a detection method that utilizes high-level semantic features of faces to identify inconsistencies in temporal domain. We introduce UniForensics, a novel deepfake detection framework that leverages a transformer-based video classification network, initialized with a meta-functional face encoder for enriched facial representation. In this way, we can take advantage of both the powerful spatio-temporal model and the high-level semantic information of faces. Furthermore, to leverage easily accessible real face data and guide the model in focusing on spatio-temporal features, we design a Dynamic Video Self-Blending (DVSB) method to efficiently generate training samples with diverse spatio-temporal forgery traces using real facial videos. Based on this, we advance our framework with a two-stage training approach: The first stage employs a novel self-supervised contrastive learning, where we encourage the network to focus on forgery traces by impelling videos generated by the same forgery process to have similar representations. On the basis of the representation learned in the first stage, the second stage involves fine-tuning on face forgery detection dataset to build a deepfake detector. Extensive experiments validates that UniForensics outperforms existing face forgery detection methods in generalization ability and robustness. In particular, our method achieves 95.3% and 77.2% cross dataset AUC on the challenging Celeb-DFv2 and DFDC respectively. Code will be made publicly available. Ziyuan Fang, Tianyi Wei, Wenbo Zhou 0004, Zhanyi Wang, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | SafeGuider: Robust and Practical Content Safety Control for Text-to-Image ModelsabstractText-to-image models have shown remarkable capabilities in generating high-quality images from natural language descriptions. However, these models are highly vulnerable to adversarial prompts, which can bypass safety measures and produce harmful content. Despite various defensive strategies, achieving robustness against attacks while maintaining practical utility in real-world applications remains a significant challenge. To address this issue, we first conduct an empirical study of the text encoder in the Stable Diffusion (SD) model, which is a widely used and representative text-to-image model. Our findings reveal that the [EOS] token acts as a semantic aggregator, exhibiting distinct distributional patterns between benign and adversarial prompts in its embedding space. Building on this insight, we introduce SafeGuider, a two-step framework designed for robust safety control without compromising generation quality. SafeGuider combines an embedding-level recognition model with a safety-aware feature erasure beam search algorithm. This integration enables the framework to maintain high-quality image generation for benign prompts while ensuring robust defense against both in-domain and out-of-domain attacks. SafeGuider demonstrates exceptional effectiveness in minimizing attack success rates, achieving a maximum rate of only 5.48% across various attack scenarios. Moreover, instead of refusing to generate or producing black images for unsafe prompts, SafeGuider generates safe and meaningful images, enhancing its practical utility. In addition, SafeGuider is not limited to the SD model and can be effectively applied to other text-to-image models, such as the Flux model, demonstrating its versatility and adaptability across different architectures. We hope that SafeGuider can shed some light on the practical deployment of secure text-to-image systems. Peigui Qi, Kunsheng Tang, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu, Tianwei Zhang 0004, Qing Guo 0005, Jie Zhang 0073 |
CCS | 3 |
| 2025 | CASAGPT: Cuboid Arrangement and Scene Assembly for Interior DesignabstractWe present a novel approach for indoor scene synthesis, which learns to arrange decomposed cuboid primitives to represent 3D objects within a scene. Unlike conventional methods that use bounding boxes to determine the placement and scale of 3D objects, our approach leverages cuboids as a straightforward yet highly effective alternative for modeling objects. This allows for compact scene generation while minimizing object intersections. Our approach, coined CasaGPT for Cuboid Arrangement and Scene Assembly, employs an autoregressive model to sequentially arrange cuboids, producing physically plausible scenes. By applying rejection sampling during the fine-tuning stage to filter out scenes with object collisions, our model further reduces intersections and enhances scene quality. Additionally, we introduce a refined dataset, 3DFRONT-NC, which eliminates significant noise presented in the original dataset, 3D-FRONT. Extensive experiments on the 3D-FRONT dataset as well as our dataset demonstrate that our approach consistently outperforms the state-of-the-art methods, enhancing the realism of generated scenes, and providing a promising direction for 3D scene synthesis. Code is available at https://github.com/CASAGPT/CASA-GPT Weitao Feng 0001, Hang Zhou 0007, Jing Liao 0001, Li Cheng 0001, Wenbo Zhou 0004 |
CVPR | 5 |
| 2025 | Segue: Side-information Guided Generative Unlearnable Examples for Facial Privacy Protection in Real WorldabstractThe widespread adoption of face recognition has raised privacy concerns regarding the collection and use of facial data. To address this, researchers have explored "unlearnable examples" by adding imperceptible perturbations during model training to prevent the model from learning target features. However, current methods are inefficient and cannot guarantee transferability and robustness at the same time, causing impracticality in the real world. To remedy it, we introduce Side-information Guided Generative Unlearnable Examples (Segue). Using a once-trained multiple-used model to generate perturbations, Segue avoids the time-consuming gradient-based approach. To improve transferability, we introduce side information such as true or pseudo labels, which are inherently consistent across different scenarios. For robustness enhancement, a distortion layer is integrated into the training pipeline. Experiments show Segue is 1000× faster than previous methods, transferable across datasets and models, and resistant to JPEG compression, adversarial training, and standard augmentations. Zhiling Zhang, Jie Zhang 0073, Wenbo Zhou 0004, Ting Xu 0004, Daiheng Gao, Zixian Guo, Qinglang Guo, Weiming Zhang 0001, Nenghai Yu |
ICASSP | 4 |
| 2025 | Beyond Sliders: Mastering the Art of Diffusion-based Image ManipulationabstractIn the realm of image generation, the quest for realism and customization has never been more pressing. While existing methods like concept sliders have made strides, they often falter when it comes to non-AIGC images, particularly images captured in real-world settings. To bridge this gap, we introduce Beyond Sliders, an innovative framework that integrates GANs and diffusion models to facilitate sophisticated image manipulation across diverse image categories. Improved upon concept sliders, our method refines the image through fine-grained guidance—both textual and visual—in an adversarial manner, leading to a marked enhancement in image quality and realism. Extensive experimental validation confirms the robustness and versatility of Beyond Sliders across a spectrum of applications. Yufei Tang, Daiheng Gao, Pingyu Wu, Wenbo Zhou 0004, Bang Zhang, Weiming Zhang 0001 |
ICME | 4 |
| 2025 | EraseAnything: Enabling Concept Erasure in Rectified Flow TransformersabstractRemoving unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching and transformer-based architectures. These advancements limit the transferability of existing concept-erasure techniques that were originally designed for the previous T2I paradigm (e.g., SD v1.4). In this work, we introduce EraseAnything, the first method specifically developed to address concept erasure within the latest flow-based T2I framework. We formulate concept erasure as a bi-level optimization problem, employing LoRA-based parameter tuning and an attention map regularizer to selectively suppress undesirable activations. Furthermore, we propose a self-contrastive learning strategy to ensure that removing unwanted concepts does not inadvertently harm performance on unrelated ones. Experimental results demonstrate that EraseAnything successfully fills the research gap left by earlier methods in this new T2I paradigm, achieving state-of-the-art performance across a wide range of concept erasure tasks. Daiheng Gao, Shilin Lu, Wenbo Zhou 0004, Jiaming Chu, Jie Zhang 0073, Mengxi Jia, Bang Zhang, Zhaoxin Fan, Weiming Zhang 0001 |
ICML | 3 |
| 2025 | T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksabstractIn recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit impressive generative capabilities, they also expose significant security risks due to their vulnerability to jailbreak attacks, where the models are manipulated to produce unsafe content such as pornography, violence, or discrimination. Existing works such as T2VSafetyBench provide preliminary benchmarks for safety evaluation, but lack systematic methods for thoroughly exploring model vulnerabilities.
To address this gap, we are the first to formalize the T2V jailbreak attack as a discrete optimization problem and propose a joint objective-based optimization framework, called \emph{T2V-OptJail}. This framework consists of two key optimization goals: bypassing the built-in safety filtering mechanisms to increase the attack success rate, preserving semantic consistency between the adversarial prompt and the unsafe input prompt, as well as between the generated video and the unsafe input prompt, to enhance content controllability. In addition, we introduce an iterative optimization strategy guided by prompt variants, where multiple semantically equivalent candidates are generated in each round, and their scores are aggregated to robustly guide the search toward optimal adversarial prompts.
We conduct large-scale experiments on several T2V models, covering both open-source models (\textit{e.g.}, Open-Sora) and real commercial closed-source models (\textit{e.g.}, Pika, Luma, Kling). The experimental results show that the proposed method improves 11.4\% and 10.0\% over the existing state-of-the-art method (SoTA) in terms of attack success rate assessed by GPT-4, attack success rate assessed by human accessors, respectively, verifying the significant advantages of the method in terms of attack effectiveness and content control. This study reveals the potential abuse risk of the semantic alignment mechanism in the current T2V model and provides a basis for the design of subsequent jailbreak defense methods. Siyuan Liang 0004, Shiqian Zhao, Rongcheng Tu, Wenbo Zhou 0004, Aishan Liu, Dacheng Tao, Siew-Kei Lam |
NeurIPS | 5 |
| 2025 | FaceTracer: Unveiling Source Identities From Swapped Face Images and Videos for Fraud PreventionabstractFace-swapping techniques have advanced rapidly with the evolution of deep learning, leading to widespread use and growing concerns about potential misuse, especially in cases of fraud. While many efforts have focused on detecting swapped face images or videos, these methods are insufficient for tracing the malicious users behind fraudulent activities. Intrusive watermark-based approaches also fail to trace unmarked identities, limiting their practical utility. To address these challenges, we introduce FaceTracer, the first non-intrusive framework specifically designed to trace the identity of the source person from swapped face images or videos. Specifically, FaceTracer leverages a disentanglement module that effectively suppresses identity information related to the target person while isolating the identity features of the source person. This allows us to extract robust identity information that can directly link the swapped face back to the original individual, aiding in uncovering the actors behind fraudulent activities. Extensive experiments demonstrate FaceTracer's effectiveness across various face-swapping techniques, successfully identifying the source person in swapped content and enabling the tracing of malicious actors involved in fraudulent activities. Additionally, FaceTracer shows strong transferability to unseen face-swapping methods including commercial applications and robustness against transmission distortions and adaptive attacks. Zhongyi Zhang 0001, Jie Zhang 0073, Wenbo Zhou 0004, Xinghui Zhou, Qing Guo 0005, Weiming Zhang 0001, Tianwei Zhang 0004, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | ADA-FInfer: Inferring Face Representations From Adaptive Select Frames for High-Visual-Quality Deepfake DetectionabstractInterpretable deepfake detection is gaining attention for providing explainable, trustworthy results, avoiding the limitations of ‘black-box’ models. Current interpretable methods focus on visible artifacts in low-visual-quality deepfakes, but these artifacts become less apparent in high-visual-quality deepfakes generated by advanced models. With advancements in deep generative models, producing high-visual-quality deepfakes has become a strategy to evade detection. To address this, we propose${\sf ADA-FInfer}$, an adaptive frame selection and interpretable face representation inference method for detecting high-visual-quality deepfakes.${\sf ADA-FInfer}$adaptively selects frames by analyzing optical flow to reveal manipulations. We also introduce an adaptive attack method that manipulates specific frames, and our adaptive selection strategy shows resistance to such attacks.${\sf ADA-FInfer}$uses an encoder to learn face representations from source and target faces, applying a representation-prediction loss to maximize the distinction between real and fake videos. To provide further insights, we employ the joint entropy, mutual information, and conditional entropy analyses to explain the method's effectiveness. Extensive experiments and ablation studies demonstrate that${\sf ADA-FInfer}$achieves promising performance in detecting high-visual-quality deepfakes. Jinwen Liang, Zheng Qin 0001, Xin Liao 0001, Wenbo Zhou 0004, Xiaodong Lin 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Compromising LLM Driven Embodied Agents With Contextual Backdoor Attacks
Aishan Liu, Yuguang Zhou, Xianglong Liu 0001, Tianyuan Zhang 0004, Siyuan Liang 0004, Jiakai Wang, Yanjun Pu, Tianlin Li, Wenbo Zhou 0004, Qing Guo 0005, Dacheng Tao |
IEEE Trans. Inf. Forensics Secur. | 10 |
| 2025 | Audio-Visual Contrastive Pre-train for Face Forgery DetectionabstractThe highly realistic avatar in the metaverse may lead to deepfakes of facial identity. Malicious users can more easily obtain the three-dimensional structure of faces, thus using deepfake technology to create counterfeit videos with higher realism. To automatically discern facial videos forged with the advancing generation techniques, deepfake detectors need to achieve stronger generalization abilities. Inspired by transfer learning, neural networks pre-trained on other large-scale face-related tasks would provide fundamental features for deepfake detection. We propose a video-level deepfake detection method based on a temporal transformer with a self-supervised audio–visual contrastive learning approach for pre-training the deepfake detector. The proposed method learns motion representations in the mouth region by encouraging the paired video and audio representations to be close while unpaired ones to be diverse. The deepfake detector adopts the pre-trained weights and partially fine-tunes on deepfake datasets. Extensive experiments show that our self-supervised pre-training method can effectively improve the accuracy and robustness of our deepfake detection model without extra human efforts. Compared with existing deepfake detection methods, our proposed method achieves better generalization ability in cross-dataset evaluations. Wenbo Zhou 0004, Dongdong Chen 0001, Weiming Zhang 0001, Ying Guo 0008, Nenghai Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | FaceRSA: RSA-Aware Facial Identity Cryptography FrameworkabstractWith the flourishing of the Internet, sharing one's photos or automated processing of faces using computer vision technology has become an everyday occurrence. While enjoying the convenience, the concern for identity privacy is also emerging. Therefore, some efforts introduced the concept of ``password'' from traditional cryptography such as RSA into the face anonymization and deanonymization task to protect the facial identity without compromising the usability of the face image. However, these methods either suffer from the poor visual quality of the synthesis results or do not possess the full cryptographic properties, resulting in compromised security. In this paper, we present the first facial identity cryptography framework with full properties analogous to RSA. Our framework leverages the powerful generative capabilities of StyleGAN to achieve megapixel-level facial identity anonymization and deanonymization. Thanks to the great semantic decoupling of StyleGAN's latent space, the identity encryption and decryption process are performed in latent space by a well-designed password mapper in the manner of editing latent code. Meanwhile, the password-related information is imperceptibly hidden in the edited latent code owing to the redundant nature of the latent space. To make our cryptographic framework possesses all the properties analogous to RSA, we propose three types of loss functions: single anonymization loss, sequential anonymization loss, and associated anonymization loss. Extensive experiments and ablation analyses demonstrate the superiority of our method in terms of the quality of synthesis results, identity-irrelevant attributes preservation, deanonymization accuracy, and completeness of properties analogous to RSA. Zhongyi Zhang 0001, Tianyi Wei, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
AAAI | 3 |
| 2024 | GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language ModelsabstractLarge language models (LLMs) have exhibited remarkable capa- bilities in natural language generation, but they have also been observed to magnify societal biases, particularly those related to gender. In response to this issue, several benchmarks have been proposed to assess gender bias in LLMs. However, these bench- marks often lack practical flexibility or inadvertently introduce biases. To address these shortcomings, we introduce GenderCARE, a comprehensive framework that encompasses innovative Criteria, bias Assessment, Reduction techniques, and Evaluation metrics for quantifying and mitigating gender bias in LLMs. To begin, we estab- lish pioneering criteria for gender equality benchmarks, spanning dimensions such as inclusivity, diversity, explainability, objectivity, robustness, and realisticity. Guided by these criteria, we construct GenderPair, a novel pair-based benchmark designed to assess gen- der bias in LLMs comprehensively. Our benchmark provides stan- dardized and realistic evaluations, including previously overlooked gender groups such as transgender and non-binary individuals. Fur- thermore, we develop effective debiasing techniques that incorpo- rate counterfactual data augmentation and specialized fine-tuning strategies to reduce gender bias in LLMs without compromising their overall performance. Extensive experiments demonstrate a significant reduction in various gender bias benchmarks, with re- ductions peaking at over 90% and averaging above 35% across 17 different LLMs. Importantly, these reductions come with minimal variability in mainstream language tasks, remaining below 2%. By offering a realistic assessment and tailored reduction of gender biases, we hope that our GenderCARE can represent a significant step towards achieving fairness and equity in LLMs. More details are available at https://github.com/kstanghere/GenderCARE-ccs24. Kunsheng Tang, Wenbo Zhou 0004, Jie Zhang 0073, Aishan Liu, Gelei Deng, Peigui Qi, Weiming Zhang 0001, Tianwei Zhang 0004, Nenghai Yu |
CCS | 2 |
| 2024 | Attribute-Aware Head Swapping Guided by 3d ModelingabstractFace manipulation has ignited the interests of both academia and industry in very recent years. Existing face manipulation methods can be roughly categorized into two types: face attribute editing and face swapping. In this paper, we focus on swapping the identity. But unlike face swapping which only changes the face region, we attempt at a more challenging task: attribute-aware head swapping. Given a source video and a target video, we replace the whole target head with the whole source head while keeping the original target attributes. To address the inherent appearance gap (e.g., hairstyle, face shape), accompanying background incompatibility and lighting difference, our method consists of three key components: 1) a generative rendering-to-real-head model for source head modeling and attribute transfer; 2) a background modeling network to fix the background incompatibility during head swapping; 3) a deep harmonization network to fix remaining issues and makes the final composited result more realistic. We compare our approach to different face manipulation methods and the experimental results demonstrate its superiority for a lot of challenging cases. Wenbo Zhou 0004, Dongdong Chen 0001, Jing Liao 0001, Jie Zhang 0073, Kejiang Chen, Weiming Zhang 0001, Nenghai Yu |
ICASSP | 1 |
| 2024 | AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRAabstractDiffusion models have achieved remarkable success in generating high-quality images. Recently, the open-source models represented by Stable Diffusion (SD) are thriving and are accessible for customization, giving rise to a vibrant community of creators and enthusiasts. However, the widespread availability of customized SD models has led to copyright concerns, like unauthorized model distribution and unconsented commercial use. To address it, recent works aim to let SD models output watermarked content for post-hoc forensics. Unfortunately, none of them can achieve the challenging white-box protection, wherein the malicious user can easily remove or replace the watermarking module to fail the subsequent verification. For this, we propose AquaLoRA as the first implementation under this scenario. Briefly, we merge watermark information into the U-Net of Stable Diffusion Models via a watermark LowRank Adaptation (LoRA) module in a two-stage manner. For watermark LoRA module, we devise a scaling matrix to achieve flexible message updates without retraining. To guarantee fidelity, we design Prior Preserving Fine-Tuning (PPFT) to ensure watermark learning with minimal impacts on model distribution, validated by proofs. Finally, we conduct extensive experiments and ablation studies to verify our design. Our code is available at github.com/Georgefwt/AquaLoRA. Weitao Feng 0001, Wenbo Zhou 0004, Jiyan He, Jie Zhang 0073, Tianyi Wei, Tianwei Zhang 0004, Weiming Zhang 0001, Nenghai Yu |
ICML | 2 |
| 2024 | Transferable Facial Privacy Protection against Blind Face Restoration via Domain-Consistent Adversarial ObfuscationabstractWith the rise of social media and the proliferation of facial recognition surveillance, concerns surrounding privacy have escalated significantly. While numerous studies have concentrated on safeguarding users against unauthorized face recognition, a new and often overlooked issue has emerged due to advances in facial restoration techniques: traditional methods of facial obfuscation may no longer provide a secure shield, as they can potentially expose anonymous information to human perception. Our empirical study shows that blind face restoration (BFR) models can restore obfuscated faces with high probability by simply retraining them on obfuscated (e.g., pixelated) faces. To address it, we propose a transferable adversarial obfuscation method for privacy protection against BFR models. Specifically, we observed a common characteristic among BFR models, namely, their capability to approximate an inverse mapping of a transformation from a high-quality image domain to a low-quality image domain. Leveraging this shared model attribute, we have developed a domain-consistent adversarial method for generating obfuscated images. In essence, our method is designed to minimize overfitting to surrogate models during the perturbation generation process, thereby enhancing the generalization of adversarial obfuscated facial images. Extensive experiments on various BFR models demonstrate the effectiveness and transferability of the proposed method. Hang Zhou 0007, Jie Zhang 0073, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
ICML | 4 |
| 2024 | Deep Image Matting With Sparse User InteractionsabstractImage matting is a fundamental and challenging problem in computer vision and graphics. Most existing matting methods leverage a user-supplied trimap as an auxiliary input to produce good alpha matte. However, obtaining high-quality trimap itself is arduous. Recently, some hint-free methods have emerged, however, the matting quality is still far behind the trimap-based methods. The main reason is that, some hints for removing semantic ambiguity and improving matting quality are essential. Apparently, there is a trade-off between interaction cost and matting quality. To balance performance and user-friendliness, we propose an improved deep image matting framework which is trimap-free and only needs sparse user click or scribble interaction to minimize the needed auxiliary constraints while still allowing interactivity. Moreover, we introduce uncertainty estimation that predicts which parts need polishing and conduct uncertainty-guided refinement. To trade off runtime against refinement quality, users can also choose different refinement modes. Experimental results show that our method performs better than existing trimap-free methods and comparably to state-of-the-art trimap-based methods with minimal user effort. Finally, we demonstrate the extensibility of our framework to video human matting without any structure modification, by adding optical flow-based sparse hint propagation and temporal consistency regularization imposed on the single frame. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Improving Deepfake Detection Generalization by Invariant Risk MinimizationabstractThe abuse of deepfake techniques has raised serious concerns about social security and ethical problems, which motivates the development of deepfake detection. However, without fully addressing the domain gap issue, existing deepfake detection methods still show weak generalization ability among datasets belonging to different domains with domain-specific characteristics like identities and generation methods, limiting their practical applications. In this paper, we propose theInvariant Domain-oriented Deepfake Detection method (ID$_{3}$), which improves the generalization of deepfake detection on multiple domains through invariant risk minimization, a novel learning paradigm that addresses the domain gap problem by jointly training a purified invariant predictor and learning an aligned invariant representation. To train a purified invariant predictor, we design theDomain Refinement Data Augmentationstrategy with self-face-swapping and region-erasing approaches, which suppresses domain-specific features and encourages the models to focus on critical domain-invariant characteristics. To learn an aligned invariant representation, we propose theDomain Calibration Batch Normalizationapproach with multiple BN branches, which normalizes input features from different domains into aligned representations during both training and testing. Extensive experiments on multiple datasets demonstrate that our framework can boost the deepfake detection generalization ability and outperform other baselines by large margins. Our codes can be found here Zixin Yin, Jiakai Wang, Yisong Xiao, Tianlin Li, Wenbo Zhou 0004, Aishan Liu, Xianglong Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | HairCLIPv2: Unifying Hair Editing via Proxy Feature BlendingabstractHair editing has made tremendous progress in recent years. Early hair editing methods use well-drawn sketches or masks to specify the editing conditions. Even though they can enable very fine-grained local control, such interaction modes are inefficient for the editing conditions that can be easily specified by language descriptions or reference images. Thanks to the recent breakthrough of cross-modal models (e.g., CLIP), HairCLIP is the first work that enables hair editing based on text descriptions or reference images. However, such text-driven and reference-driven interaction modes make HairCLIP unable to support fine-grained controls specified by sketch or mask. In this paper, we propose HairCLIPv2, aiming to support all the aforementioned interactions with one unified framework. Simultaneously, it improves upon HairCLIP with better irrelevant attributes (e.g., identity, background) preservation and unseen text descriptions support. The key idea is to convert all the hair editing tasks into hair transfer tasks, with editing conditions converted into different proxies accordingly. The editing effects are added upon the input image by blending the corresponding proxy features within the hairstyle or hair color feature spaces. Besides the unprecedented user interaction mode support, quantitative and qualitative experiments demonstrate the superiority of HairCLIPv2 in terms of editing effects, irrelevant attribute preservation and visual naturalness. Our code is available at https://github.com/wty-ustc/HairCLIPv2. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Gang Hua 0001, Nenghai Yu |
ICCV | 3 |
| 2023 | It Wasn't Me: Irregular Identity in Deepfake VideosabstractWith the rapid development in media generation technologies, the creation of DeepFake videos is within everyone’s reach. As the widespread diffusion of DeepFakes can lead to severe consequences (e.g., defamation, fake news spreading, etc.), detecting DeepFakes is becoming a crucial task within the forensic community. However, most of the existing DeepFake detectors suffer from two issues: i) they are hardly explainable as they build upon black-box data-driven techniques rather than interpretable features; ii) they are often tailored to low-level texture features, failing to generalize on low-quality DeepFake videos. In this work we propose a video DeepFake detector that aims at solving these issues. The proposed detector relies on the fact that most DeepFake generators work on a frame-by-frame basis, thus breaking the temporal consistency of facial features across frames. In particular, we noticed that facial identity features tend to be less stable in time on DeepFake videos than original ones. We therefore propose a framework trained on time series of facial identity features. The use of high-level semantic features makes the detector interpretable and robust against low-quality DeepFake videos. Extensive experiments show that our method achieves outstanding performance on low-quality DeepFake video and obtains promising results on unseen dataset evaluation. The code is available at https://github.com/HongguLiu/Identity-Inconsistency-DeepFake-Detection Honggu Liu, Paolo Bestagini, Wenbo Zhou 0004, Stefano Tubaro, Weiming Zhang 0001, Nenghai Yu |
ICIP | 4 |
| 2023 | X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionabstractCopy-Paste is a simple and effective data augmentation strategy for instance segmentation. By randomly pasting object instances onto new background images, it creates new training data for free and significantly boosts the segmentation performance, especially for rare object categories. Although diverse, high-quality object instances used in Copy-Paste result in more performance gain, previous works utilize object instances either from human-annotated instance segmentation datasets or rendered from 3D object models, and both approaches are too expensive to scale up to obtain good diversity. In this paper, we revisit Copy-Paste at scale with the power of newly emerged zero-shot recognition models (e.g., CLIP) and text2image models (e.g., StableDiffusion). We demonstrate for the first time that using a text2image model to generate images or zero-shot recognition model to filter noisily crawled images for different object categories is a feasible way to make Copy-Paste truly scalable. To make such success happen, we design a data acquisition and processing framework, dubbed ``X-Paste", upon which a systematic study is conducted. On the LVIS dataset, X-Paste provides impressive improvements over the strong baseline CenterNet2 with Swin-L as the backbone. Specifically, it archives +2.6 box AP and +2.1 mask AP gains on all classes and even more significant gains with +6.8 box AP +6.5 mask AP on long-tail classes. Dianmo Sheng, Jianmin Bao, Dongdong Chen 0001, Dong Chen 0003, Fang Wen 0001, Lu Yuan 0001, Ce Liu 0001, Wenbo Zhou 0004, Qi Chu 0001, Weiming Zhang 0001, Nenghai Yu |
ICML | 9 |
| 2023 | BiFPro: A Bidirectional Facial-data Protection Framework against DeepFakeabstractThe rapid progress of the DeepFake technique has caused severe privacy problems. Thus protecting facial data against DeepFake becomes an urgent requirement. Face protection can be regarded as a bidirectional process: Face-out-detection (FOD) and Face-in-forensics (FIF). For FOD, the detectability should be satisfied when using the protected face to replace other faces. For FIF, traceability should be guaranteed when the protected face is replaced by others. For this, we propose a Bidirectional Facial-data Protection Framework (BiFPro) to protect face data comprehensively. This framework is composed of three main parts: Watermarking embedding, Face-out-detection (FOD) and Face-in-forensics (FIF). For the FOD case, we ensure the vulnerability of the original face by embedding fragile watermarking. Once the protected facial image is used to replace other faces, the watermarking information will be corrupted in the synthesized face images which can be used to detect the authenticity of the protected facial images. As for the FIF case, we guarantee the traceability of the protected face image by embedding robust watermarking, with which the fake faces can be traced with the reserved watermarking even after the face is swapped. Experimental results demonstrate that our proposed BiFPro could generate the watermarking which is fragile to FOD and at the same time robust to FIF with an average watermark extraction success rate reaching more than 95% when defending against the four advanced DeepFake techniques. Finally, we hope this work can encourage more initiative countermeasures against DeepFake. Honggu Liu, Wenbo Zhou 0004, Han Fang 0004, Paolo Bestagini, Weiming Zhang 0001, Yuefeng Chen, Stefano Tubaro, Nenghai Yu, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 3 |
| 2023 | X-Adv: Physical Adversarial Object Attacks against X-ray Prohibited Item Detection
Aishan Liu, Jun Guo 0009, Jiakai Wang, Siyuan Liang 0004, Renshuai Tao, Wenbo Zhou 0004, Cong Liu 0006, Xianglong Liu 0001, Dacheng Tao |
USENIX Security Symposium | 6 |
| 2023 | Deepfacelab: Integrated, flexible and extensible face-swapping frameworkabstractFace swapping has drawn a lot of attention for its compelling performance. However, current deepfake methods suffer the effects of obscure workflow and poor performance. To solve these problems, we present DeepFaceLab, the current dominant deepfake framework for practical face-swapping. It provides the necessary tools as well as an easy-to-use way to conduct high-quality face-swapping. It also offers a flexible and loose coupling structure for people who need to strengthen their pipeline with other features without writing complicated boilerplate code. We detail the principles that drive the implementation of DeepFaceLab and introduce its pipeline. DeepFaceLab could achieve cinema-level results with high fidelity as our supplemental video shows. We also demonstrate the advantage of our system by comparing our approach with other face-swapping methods. Deepfake defense not only requires the research of detection but also requires the efforts of generation methods. As for a popular and practical toolkit, we encourage users to promote harmless deepfake-entertainment content on social media, reminding the public of the existence of deepfake when they are looking for entertainment. Kunlin Liu, Ivan Perov, Daiheng Gao, Nikolay Chervoniy, Wenbo Zhou 0004, Weiming Zhang 0001 |
Pattern Recognit. | 5 |
| 2023 | Coherent adversarial deepfake video generation
Honggu Liu, Wenbo Zhou 0004, Dongdong Chen 0001, Han Fang 0004, Huanyu Bian, Kunlin Liu, Weiming Zhang 0001, Nenghai Yu |
Signal Process. | 2 |
| 2022 | FInfer: Frame Inference-Based Deepfake Detection for High-Visual-Quality VideosabstractDeepfake has ignited hot research interests in both academia and industry due to its potential security threats. Many countermeasures have been proposed to mitigate such risks. Current Deepfake detection methods achieve superior performances in dealing with low-visual-quality Deepfake media which can be distinguished by the obvious visual artifacts. However, with the development of deep generative models, the realism of Deepfake media has been significantly improved and becomes tough challenging to current detection models. In this paper, we propose a frame inference-based detection framework (FInfer) to solve the problem of high-visual-quality Deepfake detection. Specifically, we first learn the referenced representations of the current and future frames’ faces. Then, the current frames’ facial representations are utilized to predict the future frames’ facial representations by using an autoregressive model. Finally, a representation-prediction loss is devised to maximize the discriminability of real videos and fake videos. We demonstrate the effectiveness of our FInfer framework through information theory analyses. The entropy and mutual information analyses indicate the correlation between the predicted representations and referenced representations in real videos is higher than that of high-visual-quality Deepfake videos. Extensive experiments demonstrate the performance of our method is promising in terms of in-dataset detection performance, detection efficiency, and cross-dataset detection performance in high-visual-quality Deepfake videos. Xin Liao 0001, Jinwen Liang, Wenbo Zhou 0004, Zheng Qin 0001 |
AAAI | 4 |
| 2022 | HairCLIP: Design Your Hair by Text and Reference ImageabstractHair editing is an interesting and challenging problem in computer vision and graphics. Many existing methods require well-drawn sketches or masks as conditional inputs for editing, however these interactions are neither straight-forward nor efficient. In order to free users from the tedious interaction process, this paper proposes a new hair editing interaction mode, which enables manipulating hair attributes individually or jointly based on the texts or reference images provided by users. For this purpose, we encode the image and text conditions in a shared embedding space and propose a unified hair editing framework by leveraging the powerful image text representation capability of the Contrastive Language-Image Pre-Training (CLIP) model. With the carefully designed network structures and loss functions, our framework can perform high-quality hair editing in a disentangled manner. Extensive experiments demonstrate the superiority of our approach in terms of manipulation accuracy, visual realism of editing results, and irrelevant attribute preservation. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Zhentao Tan, Lu Yuan 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 3 |
| 2022 | ADT: Anti-Deepfake TransformerabstractRecently almost all the mainstream deepfake detection methods use Convolutional Neural Networks (CNN) as their backbone. However, due to the overreliance on local texture information which is usually determined by forgery methods of training data, these CNN-based methods cannot generalize well to unseen data. To get out of the predicament of prior methods, in this paper, we propose a novel transformer-based framework to model both global and local information and analyze anomalies of face images. In particular, we design attention leading module, multi-forensics module and variant residual connections for deepfake detection, and leverage token-level contrast loss for more detailed supervision. Experiments on almost all popular public deepfake datasets demonstrate that our method achieves state-of-the-art performance in cross-dataset evaluation and comparable performance in intra-dataset evaluation. Ping Wang 0036, Kunlin Liu, Wenbo Zhou 0004, Hang Zhou 0007, Honggu Liu, Weiming Zhang 0001, Nenghai Yu |
ICASSP | 3 |
| 2022 | E2Style: Improve the Efficiency and Effectiveness of StyleGAN InversionabstractThis paper studies the problem of StyleGAN inversion, which plays an essential role in enabling the pretrained StyleGAN to be used for real image editing tasks. The goal of StyleGAN inversion is to find the exact latent code of the given image in the latent space of StyleGAN. This problem has a high demand for quality and efficiency. Existing optimization-based methods can produce high-quality results, but the optimization often takes a long time. On the contrary, forward-based methods are usually faster but the quality of their results is inferior. In this paper, we present a new feed-forward network "E2Style" for StyleGAN inversion, with significant improvement in terms of efficiency and effectiveness. In our inversion network, we introduce: 1) a shallower backbone with multiple efficient heads across scales; 2) multi-layer identity loss and multi-layer face parsing loss to the loss function; and 3) multi-stage refinement. Combining these designs together forms an effective and efficient method that exploits all benefits of optimization-based and forward-based methods. Quantitative and qualitative results show that our E2Style performs better than existing forward-based methods and comparably to state-of-the-art optimization-based methods while maintaining the high efficiency as well as forward-based methods. Moreover, a number of real image editing applications demonstrate the efficacy of our E2Style. Our code is available at https://github.com/wty-ustc/e2style. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Image Process. | 3 |
| 2022 | TERA: Screen-to-Camera Image Code With Transparency, Efficiency, Robustness and AdaptabilityabstractWith the rapid development of digital devices, the issue of how to transmit information among different devices with multimedia carriers has drawn much attention from the research community. This paper focuses on the important user scenario of “screen-to-camera information transmission”. Along this direction, image coding-based techniques have been shown to be the most popular and effective methods in the past decades. However, after careful study, we find that none of the existing methods can satisfy the four important properties simultaneously, i.e.,high transparency,high embedding efficiency,strong transmission robustnessandhigh adaptability to device types. This is mainly because these properties are contradictory with each other. In this paper, we thus propose a screen-to-camera image code dubbed “TERA” (transparency,efficiency,robustness andadaptability), which makes it possible to circumvent the contradiction among the above four properties for the first time. Generally, TERA adopts the color decomposition principle to ensure the visual quality and the superposition-based scheme to ensure embedding efficiency. BCH-coding-based information arrangement and a powerful attention-guided information decoding network are further designed to guarantee the robustness and adaptability. Through extensive experiments, the superiority and broad applications of our method are demonstrated. Han Fang 0004, Dongdong Chen 0001, Zehua Ma, Honggu Liu, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Multim. | 6 |
| 2022 | JPEG Robust Invertible GrayscaleabstractInvertible grayscale is a special kind of grayscale from which the original color can be recovered. Given an input color image, this seminal work tries to hide the color information into its grayscale counterpart while making it hard to recognize any anomalies. This powerful functionality is enabled by training a hiding sub-network and restoring sub-network in an end-to-end way. Despite its expressive results, two key limitations exist: 1) The restored color image often suffers from some noticeable visual artifacts in the smooth regions. 2) It is very sensitive to JPEG compression, i.e., the original color information cannot be well recovered once the intermediate grayscale image is compressed by JPEG. To overcome these two limitations, this article introduces adversarial training and JPEG simulator respectively. Specifically, two auxiliary adversarial networks are incorporated to make the intermediate grayscale images and final restored color images indistinguishable from normal grayscale and color images. And the JPEG simulator is utilized to simulate real JPEG compression during the online training so that the hiding and restoring sub-networks can automatically learn to be JPEG robust. Extensive experiments demonstrate that the proposed method is superior to the original invertible grayscale work both qualitatively and quantitatively while ensuring the JPEG robustness. We further show that the proposed framework can be applied under different types of grayscale constraints and achieve excellent results. Kunlin Liu, Dongdong Chen 0001, Jing Liao 0001, Weiming Zhang 0001, Hang Zhou 0007, Jie Zhang 0073, Wenbo Zhou 0004, Nenghai Yu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | Initiative Defense against Facial ManipulationabstractBenefiting from the development of generative adversarial networks (GAN), facial manipulation has achieved significant progress in both academia and industry recently. It inspires an increasing number of entertainment applications but also incurs severe threats to individual privacy and even political security meanwhile. To mitigate such risks, many countermeasures have been proposed. However, the great majority methods are designed in a passive manner, which is to detect whether the facial images or videos are tampered after their wide propagation. These detection-based methods have a fatal limitation, that is, they only work for ex-post forensics but can not prevent the engendering of malicious behavior. To address the limitation, in this paper, we propose a novel framework of initiative defense to degrade the performance of facial manipulation models controlled by malicious users. The basic idea is to actively inject imperceptible venom into target facial data before manipulation. To this end, we first imitate the target manipulation model with a surrogate model, and then devise a poison perturbation generator to obtain the desired venom. An alternating training strategy are further leveraged to train both the surrogate model and the perturbation generator. Two typical facial manipulation tasks: face attribute editing and face reenactment, are considered in our initiative defense framework. Extensive experiments demonstrate the effectiveness and robustness of our framework in different settings. Finally, we hope this work can shed some light on initiative countermeasures against more adversarial scenarios. Qidong Huang, Jie Zhang 0073, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
AAAI | 3 |
| 2021 | Spatial-Phase Shallow Learning: Rethinking Face Forgery Detection in Frequency DomainabstractThe remarkable success in face forgery techniques has received considerable attention in computer vision due to security concerns. We observe that up-sampling is a necessary step of most face forgery techniques, and cumulative up-sampling will result in obvious changes in the frequency domain, especially in the phase spectrum. According to the property of natural images, the phase spectrum preserves abundant frequency components that provide extra information and complement the loss of the amplitude spectrum. To this end, we present a novel Spatial-Phase Shallow Learning (SPSL) method, which combines spatial image and phase spectrum to capture the up-sampling artifacts of face forgery to improve the transferability, for face forgery detection. And we also theoretically analyze the validity of utilizing the phase spectrum. Moreover, we notice that local texture information is more crucial than high-level semantic information for the face forgery detection task. So we reduce the receptive fields by shallowing the network to suppress high-level features and focus on the local region. Extensive experiments show that SPSL can achieve the state-of-the-art performance on cross-datasets evaluation as well as multi-class classification and obtain comparable results on single dataset evaluation. Honggu Liu, Wenbo Zhou 0004, Yuefeng Chen, Yuan He 0011, Hui Xue 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 3 |
| 2021 | Improved Image Matting via Real-Time User Clicks and Uncertainty EstimationabstractImage matting is a fundamental and challenging problem in computer vision and graphics. Most existing matting methods leverage a user-supplied trimap as an auxiliary input to produce good alpha matte. However, obtaining high-quality trimap itself is arduous, thus restricting the application of these methods. Recently, some trimap-free methods have emerged, however, the matting quality is still far behind the trimap-based methods. The main reason is that, without the trimap guidance in some cases, the target network is ambiguous about which is the foreground target. In fact, choosing the foreground is a subjective procedure and depends on the user’s intention. To this end, this paper proposes an improved deep image matting framework which is trimap-free and only needs several user click interactions to eliminate the ambiguity. Moreover, we introduce a new uncertainty estimation module that can predict which parts need polishing and a following local refinement module. Based on the computation budget, users can choose how many local parts to improve with the uncertainty guidance. Quantitative and qualitative results show that our method performs better than existing trimap-free methods and comparably to state-of-the-art trimap-based methods with minimal user effort. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 3 |
| 2021 | Multi-Attentional Deepfake DetectionabstractFace forgery by deepfake is widely spread over the internet and has raised severe societal concerns. Recently, how to detect such forgery contents has become a hot research topic and many deepfake detection methods have been proposed. Most of them model deepfake detection as a vanilla binary classification problem, i.e, first use a backbone network to extract a global feature and then feed it into a binary classifier (real/fake). But since the difference between the real and fake images in this task is often subtle and local, we argue this vanilla solution is not optimal. In this paper, we instead formulate deepfake detection as a fine-grained classification problem and propose a new multi-attentional deepfake detection network. Specifically, it consists of three key components: 1) multiple spatial attention heads to make the network attend to different local parts; 2) textural feature enhancement block to zoom in the subtle artifacts in shallow features; 3) aggregate the low-level textural feature and high-level semantic features guided by the attention maps. Moreover, to address the learning difficulty of this network, we further introduce a new regional independence loss and an attention guided data augmentation strategy. Through extensive experiments on different datasets, we demonstrate the superiority of our method over the vanilla binary classifier counterparts, and achieve state-of-the-art performance. The models will be released recently at https://github.com/yoctta/multiple-attention. Wenbo Zhou 0004, Dongdong Chen 0001, Tianyi Wei, Weiming Zhang 0001, Nenghai Yu |
CVPR | 2 |
| 2021 | Deepfake Video Detection Using 3D-Attentional Inception Convolutional Neural NetworkabstractThe current spike of deepfake techniques has received considerable attention due to security concerns. To mitigate the potential risks brought by deepfake techniques, many detection methods have been proposed. However, most existing works merely leverage spatial information from separate frames and ignore valuable inter-frame temporal information. In this paper, we propose a deepfake detection scheme that uses 3D-attentional inception network. The proposed model encompasses both spatial and temporal information simultaneously with the 3D kernels. Furthermore, the channel and spatial-temporal attention modules are applied to improve detection capabilities. Comprehensive experiments demonstrate that our scheme outperforms state-of-the-art methods. Changlei Lu, Bin Liu 0016, Wenbo Zhou 0004, Qi Chu 0001, Nenghai Yu |
ICIP | 3 |
| 2021 | Adversarial defense via self-orthogonal randomization super-network
Huanyu Bian, Dongdong Chen 0001, Hang Zhou 0007, Xiaoyi Dong, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
Neurocomputing | 6 |
| 2021 | CDAE: Color decomposition-based adversarial examples for screen devices
Huanyu Bian, Hao Cui 0004, Kunlin Liu, Hang Zhou 0007, Dongdong Chen 0001, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
Inf. Sci. | 6 |
| 2021 | Adversarial batch image steganography against CNN-based pooled steganalysis
Li Li 0103, Weiming Zhang 0001, Chuan Qin 0003, Kejiang Chen, Wenbo Zhou 0004, Nenghai Yu |
Signal Process. | 5 |
| 2020 | Model Watermarking for Image Processing NetworksabstractDeep learning has achieved tremendous success in numerous industrial applications. As training a good model often needs massive high-quality data and computation resources, the learned models often have significant business values. However, these valuable deep models are exposed to a huge risk of infringements. For example, if the attacker has the full information of one target model including the network structure and weights, the model can be easily finetuned on new datasets. Even if the attacker can only access the output of the target model, he/she can still train another similar surrogate model by generating a large scale of input-output training pairs. How to protect the intellectual property of deep models is a very important but seriously under-researched problem. There are a few recent attempts at classification network protection only.In this paper, we propose the first model watermarking framework for protecting image processing models. To achieve this goal, we leverage the spatial invisible watermarking mechanism. Specifically, given a black-box target model, a unified and invisible watermark is hidden into its outputs, which can be regarded as a special task-agnostic barrier. In this way, when the attacker trains one surrogate model by using the input-output pairs of the target model, the hidden watermark will be learned and extracted afterward. To enable watermarks from binary bits to high-resolution images, both traditional and deep spatial invisible watermarking mechanism are considered. Experiments demonstrate the robustness of the proposed watermarking mechanism, which can resist surrogate models learned with different network structures and objective functions. Besides deep models, the proposed method is also easy to be extended to protect data and traditional image processing algorithms. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Han Fang 0004, Weiming Zhang 0001, Wenbo Zhou 0004, Hao Cui 0004, Nenghai Yu |
AAAI | 6 |
| 2020 | Shortening the Cover for Fast JPEG SteganographyabstractRecently, the most effective steganographic schemes for JPEG images are based on minimal distortion model with Syndrome-Trellis Codes (STCs) as the coding method. However, the execution time of STCs will be severely long for message embedding to the cover object of large size, which cannot meet the demand for real-time communication in a real-world application. According to the time complexity O(2hn), it is suggested in the STCs to accelerate the embedding process by decreasing the constraint height h. However, smaller h corresponds to lower steganographic security. In this paper, we investigate the possibility of shortening the cover (reducing the length n) for speeding up the execution of STCs without weakening the steganographic security. After introducing some properties of cover selection with proofs, we propose several algorithms designed for JPEG images to construct a preferable shortened cover containing DCT coefficients of smaller costs as much as possible. The experimental results display the superiority of the proposed algorithm on the speed profit and the security when compared with the method of decreasing h. With confidence, a JPEG image of arbitrary quality factor can be safely shortened to 1/4 of the original, and correspondingly the execution of STCs can be four times faster. Weixiang Li, Wenbo Zhou 0004, Weiming Zhang 0001, Chuan Qin 0003, Huanhuan Hu, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | DUP-Net: Denoiser and Upsampler Network for 3D Adversarial Point Clouds DefenseabstractNeural networks are vulnerable to adversarial examples, which poses a threat to their application in security sensitive systems. We propose a Denoiser and UPsampler Network (DUP-Net) structure as defenses for 3D adversarial point cloud classification, where the two modules reconstruct surface smoothness by dropping or adding points. In this paper, statistical outlier removal (SOR) and a data-driven upsampling network are considered as denoiser and upsampler respectively. Compared with baseline defenses, DUP-Net has three advantages. First, with DUP-Net as a defense, the target model is more robust to white-box adversarial attacks. Second, the statistical outlier removal provides added robustness since it is a non-differentiable denoising operation. Third, the upsampler network can be trained on a small dataset and defends well against adversarial attacks generated from other point cloud datasets. We conduct various experiments to validate that DUP-Net is very effective as defense in practice. Our best defense eliminates 83.8% of C&W and l2 loss based attack (point shifting), 50.0% of C&W and Hausdorff distance loss based attack (point adding) and 9.0% of saliency map based attack (point dropping) under 200 dropped points on PointNet. Hang Zhou 0007, Kejiang Chen, Weiming Zhang 0001, Han Fang 0004, Wenbo Zhou 0004, Nenghai Yu |
ICCV | 5 |
| 2019 | Controversial 'pixel' prior rule for JPEG adaptive steganographyabstractCurrently, the most successful model for image adaptive steganography is the framework of minimal distortion, in which a reasonable definition of costs can improve the security level. In the authors' previous work, they developed a rule for cost reassignment in spatial domain called the ‘controversial pixel prior (CPP)’ rule, which defines controversial pixels by utilizing the controversies among several comparable schemes. The CPP rule gives controversial pixels higher modification priorities. In this study, they investigate migrating the CPP rule from the spatial domain to the joint photographic experts group (JPEG) domain and name it the J‐CPP rule. In JPEG images, the cover elements are discrete cosine transform (DCT) coefficients and variant factors mayinfluence the distortion definition includingquantisation step, inter‐blocks correlation and block energy. However, there is no evidence to reveal which factor is of highest priority for promoting security. In this work, they investigate which factor is more helpful in promoting J‐CPP rule, and they finally determine to set the spatial block residual as a penalty to perfect J‐CPP rule. Through extensive experiments on different JPEG steganographic algorithms and steganalysis features, they demonstrate that the J‐CPP rule can improve the security of JPEG adaptive steganography. Wenbo Zhou 0004, Weixiang Li, Kejiang Chen, Hang Zhou 0007, Weiming Zhang 0001, Nenghai Yu |
IET Image Process. | 1 |
| 2019 | Defining Cost Functions for Adaptive JPEG Steganography at the MicroscaleabstractMinimal distortion steganography is the most successful model for adaptive steganography, in which the cost function determines the security. Texture complexity is the major factor in defining cost function in images. In this paper, we proposed a method to improve the cost function of JPEG steganography by exploiting the texture in microscale. The proposed scheme is designed by using a “microscope” to highlight details in an image, so that distortion definition can be more refined. Linear unsharp masking acts as the microscope, because it can accentuate the texture region as well as maintain the original characteristics of images. Inter-block spreading rule is proposed to further strengthen the security. We improve the state-of-the-art schemes, J-UNIWARD and UERD, as J-UNIWARD has outstanding performance on resisting detection while UERD has significant lower computational complexity. In order to keep high efficiency of UERD, filtering in the DCT domain is introduced. Extending experiments show that in most cases the proposed methods (J-MSUNIWARD and MSUERD) can achieve a higher level of security than the original methods. Kejiang Chen, Hang Zhou 0007, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Defining Joint Distortion for JPEG SteganographyabstractRecent studies have shown that the non-additive distortion model of Decomposing Joint Distortion ($DeJoin$) can work well for spatial image steganography by defining joint distortion with the principle of Synchronizing Modification Directions (SMD). However, no principles have yet produced to instruct the definition of joint distortion for JPEG steganography. Experimental results indicate that SMD can not be directly used for JPEG images, which means that simply pursuing modification directions clustered does not help improve the steganographic security. In this paper, we inspect the embedding change from the spatial domain and propose a principle of Block Boundary Continuity (BBC) for defining JPEG joint distortion, which aims to restrain blocking artifacts caused by inter-block adjacent modifications and thus effectively preserve the spatial continuity at block boundaries. According to BBC, whether inter-block adjacent modifications should be synchronized or desynchronized is related to the DCT mode and the adjacent direction of inter-block coefficients (horizontal or vertical). When built into $DeJoin$, experiments demonstrate that BBC does help improve state-of-the-art additive distortion schemes in terms of relatively large embedding payloads against modern JPEG steganalyzers. Weixiang Li, Weiming Zhang 0001, Kejiang Chen, Wenbo Zhou 0004, Nenghai Yu |
IH&MMSec | 4 |
| 2017 | A New Rule for Cost Reassignment in Adaptive SteganographyabstractIn steganography schemes, the distortion function is used to define modification costs on cover elements, which is distinctly vital to the security of modern adaptive steganography. There are several successful rules for reassigning the costs defined by a given distortion function, which can promote the security level of the corresponding steganographic algorithm. In this paper, we propose a novel cost reassignment rule, which is applied to not one but a batch of existing distortion functions. We find that the costs assigned on some pixels by several steganographic methods may be very different even though these methods exhibit close security levels. We call such pixels “controversial pixel”. Experimental results show that steganalysis features are not sensitive to controversial pixels; therefore, these pixels are suitable to carry more payloads. We name this rule the controversial pixels prior (CPP) rule. Following the rule, we propose a cost reassignment scheme. Through extensive experiments on several kinds of stego algorithms, steganalysis features, and cover databases, we demonstrate that the CPP rule can improve the security of the state-of-the-art steganographic algorithms for spatial images. Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Inf. Forensics Secur. | 1 |