VLDB 2026 Research / reviewers in the wild / expert
Jing Dong 0003
dblp:85/1692-3
· DBLP profile ↗
81ranked-venue papers
5as first author
49since 2021 · last 2026
0000-0002-2763-7832ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 2 first-author · 36 since 2021Artificial intelligence and machine learning · 30 · 1 first-author · 27 since 2021Security and privacy · 15 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dark Miner: Towards combating residuals in concept erasure for text-to-image diffusion models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 6 |
| 2026 | DREAM: A Benchmark Study for Deepfake PhotoRealism AssessMentabstractDeep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its computational modeling, imitation, is also a significant problem but lacks adequate study. In this paper, we focus on the photorealism assessment of deepfakes, which is defined as the automatic assessment of deepfake photorealism that approximates human perception of deepfakes. It is important for evaluating the quality, deceptiveness of deepfakes which can be used for predicting the influence of deepfakes on Internet, it also has potentials in improving the deepfake generation process by serving as a critic. This paper promotes this new direction by presenting a comprehensive benchmark called DREAM, which stands for Deepfake photoREalism AssessMent. It is comprised of a deepfake video dataset of diverse quality, a large scale annotation that includes 140, 000 photorealism scores, textual descriptions obtained from 3, 500 human annotators, a comprehensive evaluation, analysis of 18 representative photorealism assessment methods, including recent large vision language model based methods, a newly proposed description-aligned CLIP method. The benchmark, insights included in this study can lay the foundation for future research in this direction, other related areas. Bo Peng 0002, Zichuan Wang, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Probing unlearned diffusion models: A transferable adversarial attack perspective
Xiaoxuan Han, Wei Wang 0025, Yang Li 0255, Jing Dong 0003 |
Pattern Recognit. | 5 |
| 2025 | Unveiling Deepfakes with Latent Diffusion Counterfactual ExplanationsabstractDeepfake technology, driven by deep learning, produces highly convincing synthetic media, raising concerns about misuse. While DeepFake detection models have achieved impressive accuracy, but due to the difficulty of distinguishing fake from real, interpretability remains challenging that humans cannot understand or trust the detection results. We propose a novel approach to enhance interpretability by generating counterfactual explanations. By integrating ensemble classifier loss and text instructions into the fine-tuning of a Latent Diffusion Model, our method effectively improves the quality and efficiency of generated counterfactual explanations. Experiments on DeepFake datasets validate the effectiveness of our approach, contributing the interpretability of Deepfake detection. Bo Peng 0002, Jing Dong 0003, Xiaoyu Zhang 0002 |
ICASSP | 3 |
| 2025 | Partial Reconstruction Error for Deepfake DetectionabstractThe rapid development of deepfake technology poses a formidable challenge to personal privacy and security, underscoring the urgent need for deepfake detection. Recently, the methods based on the reconstruction error, such as DIRE and RECCE, achieve impressive performance in forgery detection. However, their performance on facial forgery datasets is relatively poor. The reconstruction process is performed on the whole images, neglecting contextual information for reconstruction. In this paper, we propose Partial Reconstruction Error to perform deepfake detection based on the reconstruction of masked regions in an image. In this way, contextual information helps to reveal the inconsistencies between the original and reconstructed regions thereby improving the detection performance. This method outperforms the best global reconstruction-based approaches on the FF++, Celeb-DF, and DiFF datasets by 4.00%, 2.83%, and 2.67%, respectively. Zheling Meng, Bo Peng 0002, Jing Dong 0003, Beilin Chu, Wei Wang 0025 |
ICASSP | 4 |
| 2025 | Unlocking A New Paradigm In Robustness For Multi-Step Facial Forgery DetectionabstractWith the rapid advancement of face forgery technologies, the quality of manipulated images has significantly improved, posing a severe threat to information security. In response, deepfake detection has emerged as an effective countermeasure against the misuse of these technologies. Sequential deepfake detection,as a specialized extension, targets face images with multi-step manipulation. However, a key challenge in this task is defending against unknown image degradation that occurs during transformation, which is not widely addressed in previous research. This paper introduces a robust detection framework named RSFDF, aimed at enhancing detection capabilities when images are subjected to degradation operations. RSFDF incorporates two critical modules:ATEM and ESCM. ATEM assists the network in focusing on important features while suppressing irrelevant information; ESCM refines the attention mechanism to increase the model’s focus on edge contours, aiding in the judgment of sequential forgeries. Experiments show that RSFDF exhibits significant improvements in robustness against unknown image degradations. Shutiao Luo, Weinan Guan, Linna Zhou, Jing Dong 0003 |
ICIP | 4 |
| 2025 | Image-level Memorization Detection via Inversion-based Inference PerturbationabstractRecent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may need to verify whether their proprietary or personal images have been memorized by the model, even in the absence of paired prompts or related metadata. We refer to this challenge as image-level memorization detection, where current methods relying on original prompts fall short. In this work, we uncover two characteristics of memorized images after perturbing the inference procedure: lower similarity of the original images and larger magnitudes of TCNP.
Building on these insights, we propose Inversion-based Inference Perturbation (IIP), a new framework for image-level memorization detection. Our approach uses unconditional DDIM inversion to derive latent codes that contain core semantic information of original images and optimizes random prompt embeddings to introduce effective perturbation. Memorized images exhibit distinct characteristics within the proposed pipeline, providing a robust basis for detection. To support this task, we construct a comprehensive setup for the image-level memorization detection, carefully curating datasets to simulate realistic memorization scenarios. Using this setup, we evaluate our IIP framework across three different memorization settings, demonstrating its state-of-the-art performance in identifying memorized images in various settings, even in the presence of data augmentation attacks. Haokun Lin, Bo Peng 0002, Zhili Liu, Yueming Lyu, Xing Zheng, Jing Dong 0003 |
ICLR | 9 |
| 2025 | Adaptive Median Smoothing: Adversarial Defense for Unlearned Text-to-Image Diffusion Models at Inference TimeabstractText-to-image (T2I) diffusion models have raised concerns about generating inappropriate content, such as "nudity". Despite efforts to erase undesirable concepts through unlearning techniques, these unlearned models remain vulnerable to adversarial inputs that can potentially regenerate such content. To safeguard unlearned models, we propose a novel inference-time defense strategy that mitigates the impact of adversarial inputs. Specifically, we first reformulate the challenge of ensuring robustness in unlearned diffusion models as a robust regression problem. Building upon the naive median smoothing for regression robustness, which employs isotropic Gaussian noise, we develop a generalized median smoothing framework that incorporates anisotropic noise. Based on this framework, we introduce a token-wise Adaptive Median Smoothing method that dynamically adjusts noise intensity according to each token’s relevance to target concepts. Furthermore, to improve inference efficiency, we explore implementations of this adaptive method at the text-encoding stage. Extensive experiments demonstrate that our approach enhances adversarial robustness while preserving model utility and inference efficiency, outperforming baseline defense techniques. Xiaoxuan Han, Wei Wang 0025, Yang Li 0255, Jing Dong 0003 |
ICML | 5 |
| 2025 | Concept Corrector: Erase Concepts on the Fly for Text-to-Image Diffusion Models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Yueming Lyu, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 6 |
| 2025 | Beyond Inserting: Learning Subject Embedding for Semantic-Fidelity Personalized Diffusion GenerationabstractText-to-Image (T2I) personalization based on advanced diffusion models (e.g., Stable Diffusion), which aims to generate images of target subjects given various prompts, has drawn huge attention. However, when users require personalized image generation for specific subjects such as themselves or their pet cat, the T2I models fail to accurately generate their subject-preserved images. The main problem is that pre-trained T2I models do not learn the T2I mapping between the target subjects and their corresponding visual contents. Even if multiple target subject images are provided, previous personalization methods either failed to accurately fit the subject region or lost the interactive generative ability with other existing concepts in T2I model space. For example, they are unable to generate T2I-aligned and semantic-fidelity images for the given prompts with other concepts such as scenes (“Eiffel Tower”), actions (“holding a basketball”), and facial attributes (“eyes closed”). In this paper, we focus on inserting accurate and interactive subject embedding into the Stable Diffusion Model for semantic-fidelity personalized generation using one image. We address this challenge from two perspectives: subject-wise attention loss and semantic-fidelity token optimization. Specifically, we propose a subject-wise attention loss to guide the subject embedding onto a manifold with high subject identity similarity and diverse interactive generative ability. Then, we optimize one subject representation as multiple per-stage tokens, and each token contains two disentangled features. This expansion of the textual conditioning space enhances the semantic control, thereby improving semantic-fidelity. We conduct extensive experiments on the most challenging subjects, face identities, to validate that our results exhibit superior subject accuracy and fine-grained manipulation ability. We further validate the generalization of our methods on various non-face subjects. Yang Li 0255, Wei Wang 0025, Jing Dong 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Noise-Informed Diffusion-Generated Image Detection With Anomaly AttentionabstractWith the rapid development of image generation technologies, especially the advancement of Diffusion Models, the quality of synthesized images has significantly improved, raising concerns among researchers about information security. To mitigate the malicious abuse of diffusion models, diffusion-generated image detection has proven to be an effective countermeasure. However, a key challenge for forgery detection is generalising to diffusion models not seen during training. In this paper, we address this problem by focusing on image noise. We observe that images from different diffusion models share similar noise patterns, distinct from genuine images. Building upon this insight, we introduce a novel Noise-Aware Self-Attention (NASA) module that focuses on noise regions to capture anomalous patterns. To implement a SOTA detection model, we incorporate NASA into Swin Transformer, forming an novel detection architecture NASA-Swin. Additionally, we employ a cross-modality fusion embedding to combine RGB and noise images, along with a channel mask strategy to enhance feature learning from both modalities. Extensive experiments demonstrate the effectiveness of our approach in enhancing detection capabilities for diffusion-generated images. When encountering unseen generation methods, our approach achieves the state-of-the-art performance. Weinan Guan, Wei Wang 0025, Bo Peng 0002, Ziwen He, Jing Dong 0003, Haonan Cheng |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Latent Watermark: Inject and Detect Watermarks in Latent Diffusion SpaceabstractWatermarking is a tool for actively identifying and attributing the images generated by latent diffusion models. Existing methods face the dilemma of image quality and watermark robustness. Watermarks with superior image quality usually have inferior robustness against attacks such as blurring and JPEG compression, while watermarks with superior robustness usually significantly damage image quality. This dilemma stems from the traditional paradigm where watermarks are injected and detected in pixel space, relying on pixel perturbation for watermark detection and resilience against attacks. In this paper, we highlight that an effective solution to the problem is to both inject and detect watermarks in the latent diffusion space, and propose Latent Watermark with a progressive training strategy. It weakens the direct connection between quality and robustness and thus alleviates their contradiction. We conduct evaluations on two datasets and against 10 watermark attacks. Six metrics measure the image quality and watermark robustness. Results show that compared to the recently proposed methods such as StableSignature, StegaStamp, RoSteALS, LaWa, TreeRing, and DiffuseTrace, LW not only surpasses them in terms of robustness but also offers superior image quality. Zheling Meng, Bo Peng 0002, Jing Dong 0003 |
IEEE Trans. Multim. | 3 |
| 2025 | Exploiting Backdoors of Face Synthesis Detection with Natural TriggersabstractDeep neural networks have enhanced face synthesis detection in discriminating Artificial Intelligence Generated Content (AIGC). However, their security is threatened by the injection of carefully crafted triggers during model training (i.e., backdoor attacks). Although existing backdoor defenses and manual data selection are able to mitigate those using human-eye-sensitive triggers, such as patches or adversarial noises, the more challenging natural backdoor triggers remain insufficiently researched. To further investigate natural triggers, we propose a novel analysis-by-synthesis backdoor attack against face synthesis detection models, which embeds natural triggers in the latent space. We study such backdoor vulnerability from two perspectives: (1) Model Discrimination (Optimization-Based Trigger) : we adopt a substitute detection model and find the trigger by minimizing the cross-entropy loss; (2) Data Distribution (Custom Trigger): we manipulate the uncommon facial attributes in the long-tailed distribution to generate poisoned samples without the supervision from detection models. Furthermore, to evaluate the detection models toward the latest AIGC, we utilize both the state-of-the-art StyleGAN and Stable Diffusion for trigger generation. Finally, these backdoor triggers introduce specific semantic features to the generated poisoned samples (e.g., skin textures and smile), which are more natural and robust. Extensive experiments show that our method is superior over existing pixel space backdoor attacks on three levels: (1) Attack Success Rate : achieving an attack success rate exceeding 99 \(\%\) , comparable to baseline methods, with less than 0.1 \(\%\) model accuracy drop and under 3 \(\%\) poisoning rate; (2) Backdoor Defense : showing superior robustness when faced with existing backdoor defenses (e.g., surpassing baseline methods by over 30 \(\%\) after a 15 \({}^{\circ}\) rotation); (3) Human Inspection : being less human-eye-sensitive from a user study with 46 participants and a collection of 2,300 data points. Xiaoxuan Han, Wei Wang 0025, Ziwen He, Jing Dong 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisabstractAudio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations. Wei Wang 0025, Bo Peng 0002, Yingya Zhang, Jing Dong 0003, Tieniu Tan |
AAAI | 6 |
| 2024 | Learning Dense Correspondence for NeRF-Based Face ReenactmentabstractFace reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment in photo-realism and 3D consistency. However, establishing dense correspondence between different face NeRFs is non-trivial, because implicit representations lack ground-truth correspondence annotations like mesh-based 3D parametric models (e.g., 3DMM) with index-aligned vertexes. Although aligning 3DMM space with NeRF-based face representations can realize motion control, it is sub-optimal for their limited face-only modeling and low identity fidelity. Therefore, we are inspired to ask: Can we learn the dense correspondence between different NeRF-based face representations without a 3D parametric model prior? To address this challenge, we propose a novel framework, which adopts tri-planes as fundamental NeRF representation and decomposes face tri-planes into three components: canonical tri-planes, identity deformations, and motion. In terms of motion control, our key contribution is proposing a Plane Dictionary (PlaneDict) module, which efficiently maps the motion conditions to a linear weighted addition of learnable orthogonal plane bases. To the best of our knowledge, our framework is the first method that achieves one-shot multi-view face reenactment without a 3D parametric model prior. Extensive experiments demonstrate that we produce better results in fine-grained motion control and identity preservation than previous methods. Wei Wang 0025, Yushi Lan, Bo Peng 0002, Jing Dong 0003 |
AAAI | 7 |
| 2024 | S3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis
Wei Wang 0025, Yifeng Ma 0001, Bo Peng 0002, Yingya Zhang, Jing Dong 0003 |
ECCV (10) | 7 |
| 2024 | Counterfactual Explanations for Face Forgery Detection via Adversarial Removal of ArtifactsabstractHighly realistic AI generated face forgeries known as deepfakes have raised serious social concerns. Although DNN-based face forgery detection models have achieved good performance, they are vulnerable to latest generative methods that have less forgery traces and adversarial attacks. This limitation of generalization and robustness hinders the credibility of detection results and requires more explanations. In this work, we provide counterfactual explanations for face forgery detection from an artifact removal perspective. Specifically, we first invert the forgery images into the StyleGAN latent space, and then adversarially optimize their latent representations with the discrimination supervision from the target detection model. We verify the effectiveness of the proposed explanations from two aspects: (1) Counterfactual Trace Visualization: the enhanced forgery images are useful to reveal artifacts by visually contrasting the original images and two different visualization methods; (2) Transferable Adversarial Attacks: the adversarial forgery images generated by attacking the detection model are able to mislead other detection models, implying the removed artifacts are general. Extensive experiments demonstrate that our method achieves over 90% attack success rate and superior attack transferability. Compared with naive adversarial noise methods, our method adopts both generative and discriminative model priors, and optimize the latent representations in a synthesis-by-analysis way, which forces the search of counterfactual explanations on the natural face manifold. Thus, more general counterfactual traces can be found and better adversarial attack transferability can be achieved. Our code is available at https://github.com/yangli-lab/Artifact-Eraser/. Yang Li 0255, Wei Wang 0025, Ziwen He, Bo Peng 0002, Jing Dong 0003 |
ICME | 6 |
| 2024 | SPI2I: Structure-Preserved Image-to-Image Translation with Diffusion Models
Beibei Dong, Bo Peng 0002, Jing Dong 0003 |
ICPR (21) | 3 |
| 2024 | Freestyle 3D-Aware Portrait Synthesis Based on Compositional Generative Priors
Tianxiang Ma, Jianxin Sun 0003, Yingya Zhang, Jing Dong 0003 |
ICPR (6) | 5 |
| 2024 | Mitigating Social Biases in Text-to-Image Diffusion Models via Linguistic-Aligned Attention GuidanceabstractRecent advancements in text-to-image generative models have showcased remarkable capabilities across various tasks. However, these powerful models have revealed the inherent risks of social biases. Such biases can propagate distorted real-world perspectives and spread unforeseen prejudice and discrimination. Current debiasing methods are primarily designed for scenarios with a single individual in the image and exhibit homogenous race or gender when multiple individuals are involved, harming the diversity of social groups within the image. To address this problem, we consider the semantic consistency between text prompts and generated images in text-to-image diffusion models to identify how biases are generated. We propose a novel method to locate where the biases are based on different tokens and then mitigate them for each individual. Specifically, we introduce a Linguistic-aligned Attention Guidance module consisting of Block Voting and Linguistic Alignment, to effectively locate the semantic regions related to biases. Additionally, we employ Fair Inference in these regions to generate fair attributes across arbitrary distributions while preserving the original structural and semantic information. Extensive experiments and analyses demonstrate our method outperforms existing methods for debiasing with multiple individuals across various scenarios. Yueming Lyu, Ziwen He, Bo Peng 0002, Jing Dong 0003 |
ACM Multimedia | 5 |
| 2024 | ST-SBV: Spatial-Temporal Self-Blended Videos for Deepfake Detection
Weinan Guan, Wei Wang 0025, Bo Peng 0002, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 4 |
| 2024 | Artifact feature purification for cross-domain detection of AI-generated images
Zheling Meng, Bo Peng 0002, Jing Dong 0003, Tieniu Tan, Haonan Cheng |
Comput. Vis. Image Underst. | 3 |
| 2024 | InfoStyler: Disentanglement Information Bottleneck for Artistic Style TransferabstractArtistic style transfer aims to transfer the style of an artwork to a photograph while maintaining its original overall content. Many prior works focus on designing various transfer modules to transfer the style statistics to the content image. Although effective, ignoring the clear disentanglement of the content features and the style features from the first beginning, they have difficulty in balancing between content preservation and style transferring. To tackle this problem, we propose a novel information disentanglement method, named InfoStyler, to capture the minimal sufficient information for both content and style representations from the pre-trained encoding network. InfoStyler formulates the disentanglement representation learning as an information compression problem by eliminating style statistics from the content image and removing the content structure from the style image. Besides, to further facilitate disentanglement learning, a cross-domain Information Bottleneck (IB) learning strategy is proposed by reconstructing the content and style domains. Extensive experiments demonstrate that our InfoStyler can synthesize high-quality stylized images while balancing content structure preservation and style pattern richness. Yueming Lyu, Bo Peng 0002, Jing Dong 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Improving Generalization of Deepfake Detectors by Imposing Gradient RegularizationabstractThe rapid development of face forgery technology has posed a significant threat to information security. While deepfake detection has proven to be an effective countermeasure, it often struggles to detect fake images generated by unknown forgery methods. Thus, the generalization ability of deepfake detectors to unseen forgery data is a critical concern. Despite many efforts aimed at discovering new forgery artifacts, they often fail to generalize to new manipulation technologies. In this paper, we tackle this challenge by focusing on the difference in texture patterns between training forgeries and unseen forgeries, which can lead to a degradation of generalization. Based on this principle, we propose a new conjecture that encourages deepfake detectors to reduce their sensitivity to forgery texture patterns, thereby improving the detection performance. To this end, we introduce an additional gradient regularization term to the original empirical loss during training. However, computing the Hessian matrix in the gradient calculation process of the regularization term poses a computational complexity. In order to overcome this issue, we optimize the formulation of the gradient regularization term using a first-order approximation method based on Taylor expansion and design a Perturbation Injection Module (PIM) to simplify the implementation process. Additionally, we provide a theoretical analysis from an optimization perspective and explore an interesting aspect of our method. Extensive experiments demonstrate the effectiveness of our approach in improving the generalization ability of deepfake detectors. Importantly, our method is orthogonal to recent advancements in powerful backbones and training data augmentation techniques. When combined with other effective techniques, our method achieves state-of-the-art experimental results. Weinan Guan, Wei Wang 0025, Jing Dong 0003, Bo Peng 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | DRAN: Detailed Region-Adaptive Normalization for Conditional Image SynthesisabstractIn recent years, conditional image synthesis has attracted growing attention due to its controllability in the image generation process. Although recent works have achieved realistic results, most of them have difficulty handling fine-grained styles with subtle details. To address this problem, a novel normalization module, named Detailed Region-Adaptive Normalization (DRAN), is proposed. It adaptively learns both fine-grained and coarse-grained style representations. Specifically, we first introduce a multi-level structure, Spatiality-aware Pyramid Pooling, to guide the model to learn coarse-to-fine features. Then, to adaptively fuse different levels of styles, we propose Dynamic Gating, making it possible to adaptively fuse different levels of styles according to different spatial regions. Finally, we collect a new makeup dataset (Makeup-Complex dataset) that contains a wide range of complex makeup styles with diverse poses and expressions. To evaluate the effectiveness and show the general use of our method, we conduct a set of experiments on makeup transfer and semantic image synthesis. Quantitative and qualitative experiments show that equipped with DRAN, simple baseline models are able to achieve promising improvements in complex style transfer and detailed texture synthesis. Yueming Lyu, Peibin Chen, Jingna Sun, Bo Peng 0002, Jing Dong 0003 |
IEEE Trans. Multim. | 6 |
| 2023 | Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance FieldabstractRecently 3D-aware GAN methods with neural radiance field have developed rapidly. However, current methods model the whole image as an overall neural radiance field, which limits the partial semantic editability of synthetic results. Since NeRF renders an image pixel by pixel, it is possible to split NeRF in the spatial dimension. We propose a Compositional Neural Radiance Field (CNeRF) for semantic 3D-aware portrait synthesis and manipulation. CNeRF divides the image by semantic regions and learns an independent neural radiance field for each region, and finally fuses them and renders the complete image. Thus we can manipulate the synthesized semantic regions independently, while fixing the other parts unchanged. Furthermore, CNeRF is also designed to decouple shape and texture within each semantic region. Compared to state-of-the-art 3D-aware GAN methods, our approach enables fine-grained semantic region manipulation, while maintaining high-quality 3D-consistent synthesis. The ablation studies show the effectiveness of the structure and loss function used by our method. In addition real image inversion and cartoon portrait 3D editing experiments demonstrate the application potential of our method. Tianxiang Ma, Bingchuan Li, Jing Dong 0003, Tieniu Tan |
AAAI | 4 |
| 2023 | CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image TranslationabstractExemplar-based image translation refers to the task of generating images with the desired style, while conditioning on certain input image. Most of the current methods learn the correspondence between two input domains and lack the mining of information within the domain. In this paper, we propose a more general learning approach by considering two domain features as a whole and learning both inter-domain correspondence and intra-domain potential information interactions. Specifically, we propose a Cross-domain Feature Fusion Transformer (CFFT) to learn inter- and intra-domain feature fusion. Based on CFFT, the proposed CFFT-GAN works well on exemplar-based image translation. Moreover, CFFT-GAN is able to decouple and fuse features from multiple domains by cascading CFFT modules. We conduct rich quantitative and qualitative experiments on several image translation tasks, and the results demonstrate the superiority of our approach compared to state-of-the-art methods. Ablation studies show the importance of our proposed CFFT. Application experimental results reflect the potential of our method. Tianxiang Ma, Bingchuan Li, Wei Liu 0035, Miao Hua, Jing Dong 0003, Tieniu Tan |
AAAI | 5 |
| 2023 | Designing A 3d-Aware Stylenerf Encoder for Face EditingabstractGAN inversion has been exploited in many face manipulation tasks, but 2D GANs often fail to generate multi-view 3D consistent images. The encoders designed for 2D GANs are not able to provide sufficient 3D information for the inversion and editing. Therefore, 3D-aware GAN inversion is proposed to increase the 3D editing capability of GANs. However, the 3D-aware GAN inversion remains under-explored. To tackle this problem, we propose a 3D-aware (3Da) encoder for GAN inversion and face editing based on the powerful StyleNeRF model. Our proposed 3Da encoder combines a parametric 3D face model with a learnable detail representation model to generate geometry, texture and view direction codes. For more flexible face manipulation, we then design a dual-branch StyleFlow module to transfer the StyleNeRF codes with disentangled geometry and texture flows. Extensive experiments demonstrate that we realize 3D consistent face manipulation in both facial attribute editing and texture transfer. Furthermore, for video editing, we make the sequence of frame codes share a common canonical manifold, which improves the temporal consistency of the edited attributes. Wei Wang 0025, Bo Peng 0002, Jing Dong 0003 |
ICASSP | 4 |
| 2023 | DFGC-VRA: DeepFake Game Competition on Visual Realism AssessmentabstractThis paper presents the summary report on the DeepFake Game Competition on Visual Realism Assessment (DFGC-VRA). Deep-learning based face-swap videos, also known as deepfakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. There is a ongoing deepfake game between its creators and detectors, with the human in the loop. The research community has been focusing on the automatic detection of these fake videos, but the assessment of their visual realism, as perceived by human eyes, is still an unexplored dimension. Visual realism assessment, or VRA, is essential for assessing the potential impact that may be brought by a specific face-swap video, and it is also useful as a quality metric to compare different face-swap methods. This is the third edition of DFGC competitions, which focuses on the new visual realism assessment topic, different from previous ones that compete creators versus detectors. With this competition, we conduct a comprehensive study of the SOTA performance on the new task. We also release our MindSpore codes to further facilitate research in this field (https://github.com/bomb2peng/DFGC-VRA-benckmark). Bo Peng 0002, Xianyun Sun, Caiyong Wang, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Rongyu Zhang, Heng Cong, Lingzhi Fu, Yusheng Zhang, Boyuan Liu, Luka Dragar, Borut Batagelj, Peter Peer, Vitomir Struc, Xinghui Zhou, Kunlin Liu, Wenxiu Diao |
IJCB | 5 |
| 2023 | Visual Realism Assessment for Face-Swap Videos
Xianyun Sun, Beibei Dong, Caiyong Wang, Bo Peng 0002, Jing Dong 0003 |
ICIG (1) | 5 |
| 2023 | Context-Aware Talking-Head Video EditingabstractTalking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head video clips which simultaneously have accurate lip synchronization and motion smoothness. Previous approaches, including 3DMM-based (3D Morphable Model) methods and NeRF-based (Neural Radiance Field) methods, are sub-optimal in that they either require minutes of source videos and days of training time or lack the disentangled control of verbal (e.g., lip motion) and non-verbal (e.g., head pose and expression) representations for video clip insertion. In this work, we fully utilize the video context to design a novel framework for talking-head video editing, which achieves efficiency, disentangled motion control, and sequential smoothness. Specifically, we decompose this framework to motion prediction and motion-conditioned rendering: (1) We first design an animation prediction module that efficiently obtains smooth and lip-sync motion sequences conditioned on the driven speech. This module adopts a non-autoregressive network to obtain context prior and improve the prediction efficiency, and it learns a speech-animation mapping prior with better generalization to novel speech from a multi-identity video dataset. (2) We then introduce a neural rendering module to synthesize the photo-realistic and full-head video frames given the predicted motion sequence. This module adopts a pre-trained head topology and uses only few frames for efficient fine-tuning to obtain a person-specific rendering model. Extensive experiments demonstrate that our method efficiently achieves smoother editing results with higher image quality and lip accuracy using less data than previous methods. Wei Wang 0025, Jun Ling, Bo Peng 0002, Xu Tan 0003, Jing Dong 0003 |
ACM Multimedia | 6 |
| 2023 | 3D-Aware Adversarial Makeup Generation for Facial Privacy ProtectionabstractThe privacy and security of face data on social media are facing unprecedented challenges as it is vulnerable to unauthorized access and identification. A common practice for solving this problem is to modify the original data so that it could be protected from being recognized by malicious face recognition (FR) systems. However, such "adversarial examples" obtained by existing methods usually suffer from low transferability and poor image quality, which severely limits the application of these methods in real-world scenarios. In this paper, we propose a 3D-Aware Adversarial Makeup Generation GAN (3DAM-GAN). which aims to improve the quality and transferability of synthetic makeup for identity information concealing. Specifically, a UV-based generator consisting of a novel Makeup Adjustment Module (MAM) and Makeup Transfer Module (MTM) is designed to render realistic and robust makeup with the aid of symmetric characteristics of human faces. Moreover, a makeup attack mechanism with an ensemble training strategy is proposed to boost the transferability of black-box models. Extensive experiment results on several benchmark datasets demonstrate that 3DAM-GAN could effectively protect faces against various FR models, including both publicly available state-of-the-art models and commercial face verification APIs, such as Face++, Baidu, and Aliyun. Yueming Lyu, Ziwen He, Bo Peng 0002, Yunfan Liu 0001, Jing Dong 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Temporal sparse adversarial attack on sequence-based gait recognition
Ziwen He, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Pattern Recognit. | 3 |
| 2023 | AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and LocalizationabstractThe point process is a solid framework to model sequential data, such as videos, by exploring the underlying relevance. As a challenging problem for high-level video understanding, weakly supervised action recognition and localization in untrimmed videos have attracted intensive research attention. Knowledge transfer by leveraging the publicly available trimmed videos as external guidance is a promising attempt to make up for the coarse-grained video-level annotation and improve the generalization performance. However, unconstrained knowledge transfer may bring about irrelevant noise and jeopardize the learning model. This article proposes a novel adaptability decomposing encoder-decoder network to transfer reliable knowledge between the trimmed and untrimmed videos for action recognition and localization by bidirectional point process modeling, given only video-level annotations. By decomposing the original features into the domain-adaptable and domain-specific ones based on their adaptability, trimmed-untrimmed knowledge transfer can be safely confined within a more coherent subspace. An encoder-decoder-based structure is carefully designed and jointly optimized to facilitate effective action classification and temporal localization. Extensive experiments are conducted on two benchmark data sets (i.e., THUMOS14 and ActivityNet1.3), and the experimental results clearly corroborate the efficacy of our method. Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035, Jing Dong 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | DFGC 2022: The Second DeepFake Game CompetitionabstractThis paper presents the summary report on our DFGC 2022 competition. The DeepFake is rapidly evolving, and realistic face-swaps are becoming more deceptive and difficult to detect. On the other hand, methods for detecting DeepFakes are also improving. There is a two-party game between DeepFake creators and defenders. This competition provides a common platform for benchmarking the game between the current state-of-the-arts in Deep-Fake creation and detection methods. The main research question to be answered by this competition is the current state of the two adversaries when competed with each other. This is the second edition after the last year's DFGC 2021, with a new, more diverse video dataset, a more realistic game setting, and more reasonable evaluation metrics. With this competition, we aim to stimulate research ideas for building better defenses against the DeepFake threats. We also release our DFGC 2022 dataset contributed by both our participants and ourselves to enrich the DeepFake data resources for the research community (https://github.com/NiCE-X/DFGC-2022). Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Zhen Lei 0001, Siwei Lyu |
IJCB | 5 |
| 2022 | Defending Against Deepfakes with Ensemble Adversarial PerturbationabstractMaliciously manipulated images and videos, represented by prevalent deepfakes, can easily deceive human and mislead the public opinions. A great deal of effort was spent on detecting these fake images or videos. However, these detection methods always encounter various problems in practical applications. Do we have other ways to block the spread of fake image or videos? This motivates us to focus on an emerging interesting topic, disruption of deepfake generation. We propose the ensemble attacks of various types of deepfake models including facial attribute editing, face swapping and face reenactment models. With the help of hard model mining, we boost the attack success rate significantly comparing with the straightforward average ensemble. Extensive experiments demonstrate the proposed approach can successfully disrupt multiple deepfake models simultaneously under white-box or gray-box attack protocols. Weinan Guan, Ziwen He, Wei Wang 0025, Jing Dong 0003, Bo Peng 0002 |
ICPR | 4 |
| 2022 | Contrastive Knowledge Transfer for Deepfake Detection with Limited DataabstractNowadays forensics methods have shown remarkable progress in detecting maliciously crafted fake images. However, without exception, the training process of deepfake detection models requires a large number of facial images. These models are usually unsuitable for real world applications because of their overlarge size and inferiority in speed. Thus, performing data-efficient deepfake detection is of great importance. In this paper, we propose a contrastive distillation method that maximizes the lower bound of mutual information between the teacher and the student to further improve student’s accuracy in a data-limited setting. We observe that models performing deepfake detection, different from other image classification tasks, have shown high robustness when there is a drop in data amount. The proposed knowledge transfer approach is of superior performance compared with vanilla few samples training baseline and other SOTA knowledge transfer methods. We believe we are the first to perform few-sample knowledge distillation on deepfake detection. Wenqi Zhuo, Wei Wang 0025, Jing Dong 0003 |
ICPR | 4 |
| 2022 | AdaDeId: Adjust Your Identity Attribute FreelyabstractFace de-identification has drawn increasing attention in recent years. It is important to protect people’s identity information meanwhile keeping the utility of the face data in many computer vision tasks. We propose a Adaptive De-identification (AdaDeId) method, a novel approach that can freely manipulate the identity attributes of given faces. We introduce an identity decoupling representation learning method, which is based on the autoencoder decoupling model as well as our proposed Identity Decoupling Representation (IDR) loss and Content Retention (CR) loss. Our method encodes the identity information of a face into a unit spherical space, where we can continuously manipulate the identity representation vector. Various de-identified faces derived from an original face can be generated through our method and maintain high similarity to the original image contents. Quantitative and qualitative experiments demonstrate our method achieves state-of-the-art on visual quality and de-identification validity. Tianxiang Ma, Wei Wang 0025, Jing Dong 0003 |
ICPR | 4 |
| 2022 | DesignerGAN: Sketch Your Own PhotoabstractPerson image generation is a challenging problem due to the complexity of human body structure and the richness of clothing texture. Recent works have made great progress on pose transfer by using keypoints, but cannot characterize the personalized shape attributes. Hence, they have limited person image editing ability, especially in respect of shape editing. In this paper, we propose to use sketches as the expression of the target image, which can not only represent the pose and shape simultaneously but is also flexible to manipulate at the semantic level. We propose DesignerGAN, a novel two-stage model for pose transfer and shape-related attributes editing. The first stage predicts the target semantic parsing using the target sketch and obtains parsing feature maps. In the second stage, with the parsing feature maps and the scaled target sketch, we devise a domain-matching spatially-adaptive normalization method to guide target image generation in multi-level. Qualitative and quantitative comparison results demonstrate our method’s superiority over state-of-the-arts on pose transfer. Besides, we achieve flexible person image editing through simple hand-drawings on sketches. Binghao Zhao, Tianxiang Ma, Bo Peng 0002, Jing Dong 0003 |
ICPR | 4 |
| 2022 | Defeating DeepFakes via Adversarial Visual ReconstructionabstractExisting DeepFake detection methods focus on passive detection, i.e., they detect fake face images by exploiting the artifacts produced during DeepFake manipulation. These detection-based methods have their limitation that they only work for ex-post forensics but cannot erase the negative influences of DeepFakes. In this work, we propose a proactive framework for combating DeepFake before the data manipulations. The key idea is to find a well defined substitute latent representation to reconstruct target facial data, leading the reconstructed face to disable the DeepFake generation. To this end, we invert face images into latent codes with a well trained auto-encoder, and search the adversarial face embeddings in their neighbor with the gradient descent method. Extensive experiments on three typical DeepFake manipulation methods, facial attribute editing, face expression manipulation, and face swapping, have demonstrated the effectiveness of our method in different settings. Ziwen He, Wei Wang 0025, Weinan Guan, Jing Dong 0003, Tieniu Tan |
ACM Multimedia | 4 |
| 2022 | Counterfactual Image Enhancement for Explanation of Face Swap Deepfakes
Bo Peng 0002, Siwei Lyu, Wei Wang 0025, Jing Dong 0003 |
PRCV (2) | 4 |
| 2022 | Revisiting ensemble adversarial attack
Ziwen He, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Signal Process. Image Commun. | 3 |
| 2022 | A Unified Framework for High Fidelity Face Swap and Expression ReenactmentabstractFace manipulation techniques improve fast with the development of powerful image generation models. Two particular face manipulation methods, namely face swap and expression reenactment attract much attention for their flexibility and ease to generate high quality synthesis results. Recently, these two subjects are actively studied. However, most existing methods treat the two tasks separately, ignoring their underlying similarity. In this paper, we propose to tackle the two problems within a unified framework that achieves high quality synthesis results. The enabling component for our unified framework is the clean disentanglement of 3D pose, shape, and expression factors and then recombining them for different tasks accordingly. We then use the same set of 2D representations for face swap and expression reenactment tasks that are input to a common image translation model to directly generate the final synthetic images. Once trained, the proposed model can accomplish both face swap and expression reenactment tasks for previously unseen subjects. Comprehensive experiments and comparisons show that the proposed method achieves high fidelity results in multiple aspects, and it is especially good at faithfully preserving source facial shape in the face swap task, and accurately transferring facial movements in the expression reenactment task. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Exploring Adversarial Fake Images on Face ManifoldabstractImages synthesized by powerful generative adversarial network (GAN) based methods have drawn moral and privacy concerns. Although image forensic models have reached great performance in detecting fake images from real ones, these models can be easily fooled with a simple adversarial attack. But, the noise adding adversarial samples are also arousing suspicion. In this paper, instead of adding adversarial noise, we optimally search adversarial points on face manifold to generate anti-forensic fake face images. We iteratively do a gradient-descent with each small step in the latent space of a generative model, e.g. Style-GAN, to find an adversarial latent vector, which is similar to norm-based adversarial attack but in latent space. Then, the generated fake images driven by the adversarial latent vectors with the help of GANs can defeat main-stream forensic models. For examples, they make the accuracy of deepfake detection models based on Xception or EfficientNet drop from over 90% to nearly 0%, mean-while maintaining high visual quality. In addition, we find manipulating noise vectors n at different levels have different impacts on attack success rate, and the generated adversarial images mainly have changes on facial texture or face attributes. Wei Wang 0025, Hongxing Fan, Jing Dong 0003 |
CVPR | 4 |
| 2021 | MUST-GAN: Multi-Level Statistics Transfer for Self-Driven Person Image GenerationabstractPose-guided person image generation usually involves using paired source-target images to supervise the training, which significantly increases the data preparation effort and limits the application of the models. To deal with this problem, we propose a novel multi-level statistics transfer model, which disentangles and transfers multi-level appearance features from person images and merges them with pose features to reconstruct the source person images themselves. So that the source images can be used as supervision for self-driven person image generation. Specifically, our model extracts multi-level features from the appearance encoder and learns the optimal appearance representation through attention mechanism and attributes statistics. Then we transfer them to a pose-guided generator for re-fusion of appearance and pose. Our approach allows for flexible manipulation of person appearance and pose properties to perform pose transfer and clothes style transfer tasks. Experimental results on the DeepFashion dataset demonstrate our method’s superiority compared with state-of-the-art supervised and unsupervised methods. In addition, our approach also performs well in the wild. Tianxiang Ma, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003 |
CVPR | 4 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 4 |
| 2021 | SOGAN: 3D-Aware Shadow and Occlusion Robust GAN for Makeup TransferabstractIn recent years, virtual makeup applications have become more and more popular. However, it is still challenging to propose a robust makeup transfer method in the real-world environment. Current makeup transfer methods mostly work well on good-conditioned clean makeup images, but transferring makeup that exhibits shadow and occlusion is not satisfying. To alleviate it, we propose a novel makeup transfer method, called 3D-Aware Shadow and Occlusion Robust GAN (SOGAN). Given the source and the reference faces, we first fit a 3D face model and then disentangle the faces into shape and texture. In the texture branch, we map the texture to the UV space and design a UV texture generator to transfer the makeup. Since human faces are symmetrical in the UV space, we can conveniently remove the undesired shadow and occlusion from the reference image by carefully designing a Flip Attention Module (FAM). After obtaining cleaner makeup features from the reference image, a Makeup Transfer Module (MTM) is introduced to perform accurate makeup transfer. The qualitative and quantitative experiments demonstrate that our SOGAN not only achieves superior results in shadow and occlusion situations but also performs well in large pose and expression variations. Yueming Lyu, Jing Dong 0003, Bo Peng 0002, Wei Wang 0025, Tieniu Tan |
ACM Multimedia | 2 |
| 2021 | SAPS: Self-Attentive Pathway Search for weakly-supervised action localization with background-action augmentation
Xiaoyu Zhang 0002, Yaru Zhang, Haichao Shi, Jing Dong 0003 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Learning pose-invariant 3D object reconstruction from single-view images
Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 3 |
| 2019 | An Accurate LSTM Based Video Heart Rate Estimation Method
Mingyun Bian, Bo Peng 0002, Wei Wang 0025, Jing Dong 0003 |
PRCV (3) | 4 |
| 2019 | Flexible Lossy Compression for Selective Encrypted Image With Image InpaintingabstractIn this paper, a novel lossy compression scheme for encrypted image based on image inpainting is proposed. In order to maintain confidentiality, the content owner encrypts the original image through a modulo-256 addition encryption and block permutation to mask image content. Then, the third party, such as a cloud server, can compress the selective encrypted image before transmitting to the receiver. During compression, encrypted blocks are categorized into four sets corresponding to different complexity degrees in plaintext domain without the loss of security. By allocating various bit rates to the encrypted blocks from different sets, flexible compression can be achieved with difference quantization. After parsing and decoding the compressed bit stream, the receiver first recovers partial encrypted pixels and then decrypts them. The other missing pixels are further recovered with the assistance of image inpainting based on a total variation model, and the final reconstructed image can be produced. Experimental results demonstrate that the proposed scheme achieves better rate-distortion performance than some of the state-of-the-art schemes. Chuan Qin 0001, Jing Dong 0003, Xinpeng Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm RetrievalabstractThere are great demands for automatically regulating inappropriate appearance of shocking firearm images in social media or identifying firearm types in forensics. Image retrieval techniques have great potential to solve these problems. To facilitate research in this area, we introduce Firearm 14k, a large dataset consisting of over 14,000 images in 167 categories. It can be used for both fine-grained recognition and retrieval of firearm images. Recent advances in image retrieval are mainly driven by fine-tuning state-of-the-art convolutional neural networks for retrieval task. The conventional single margin contrastive loss, known for its simplicity and good performance, has been widely used. We find that it performs poorly on the Firearm 14k dataset due to: (1) Loss contributed by positive and negative image pairs is unbalanced during training process. (2) A huge domain gap exists between this dataset and ImageNet. We propose to deal with the unbalanced loss by employing a double margin contrastive loss. We tackle the domain gap issue with a two-stage training strategy, where we first fine-tune the network for classification, and then fine-tune it for retrieval. Experimental results show that our approach outperforms the conventional single margin approach by a large margin (up to 88.5% relative improvement) and even surpasses the strong triplet-loss-based approach. Jiedong Hao, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICPR | 2 |
| 2018 | Ensemble Reversible Data HidingabstractThe conventional reversible data hiding (RDH) algorithms often consider the host as a whole to embed a secret payload. In order to achieve satisfactory rate-distortion performance, the secret bits are embedded into the noise-like component of the host such as prediction errors. From the rate-distortion optimization view, it may be not optimal since the data embedding units use the identical parameters. This motivates us to present a segmented data embedding strategy for efficient RDH in this paper, in which the raw host could be partitioned into multiple subhosts such that each one can freely optimize and use the data embedding parameters. Moreover, it enables us to apply different RDH algorithms within different subhosts, which is defined as ensemble. Notice that, the ensemble defined here is different from that in machine learning. Accordingly, the conventional operation corresponds to a special case of the proposed work. Since it is a general strategy, we combine some state-of-the-art algorithms to construct a new system using the proposed embedding strategy to evaluate the rate-distortion performance. Experimental results have shown that, the ensemble RDH system could outperform the original versions in most cases, which has shown the superiority and applicability. Hanzhou Wu, Wei Wang 0025, Jing Dong 0003, Hongxia Wang 0001 |
ICPR | 3 |
| 2018 | Reversible data hiding in encrypted image with separable capability and high embedding capacity
Chuan Qin 0001, Zhihong He, Xiangyang Luo 0001, Jing Dong 0003 |
Inf. Sci. | 4 |
| 2018 | Deep-MATEM: TEM query image based cross-modal retrieval for material science literature
Qingxiao Guan, Jing Dong 0003 |
Multim. Tools Appl. | 4 |
| 2018 | Feature learning for steganalysis using convolutional neural networks
Yinlong Qian, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
Multim. Tools Appl. | 2 |
| 2018 | Image Forensics Based on Planar Contact Constraints of 3D ObjectsabstractStanding objects on planar surfaces are common to see in images, e.g., people on the ground. For most objects to stay stable on the plane, planar contact is a necessary requirement. However, 2D image splicing usually disregards this physical constraint of 3D world, leading to a potential artifact of object not attached to the plane. This paper is the first attempt to use the contact constraint of standing objects as a new clue for image forensics. Accordingly, we propose a novel approach to first reconstruct the 3D poses of standing objects and their supporting plane and then measure the contact conditions for splicing detection. To tackle the problem of unknown object shape for pose estimation, we effectively employ the prior knowledge of 3D morphable model to simultaneously estimate both shape and pose parameters by fitting to image observations. The 3D normal orientation of the supporting plane is estimated given its vanishing line. Dealing with uncertainty factors in estimations, we approximate a distribution of estimates using sampling strategies and then make the final decision. Particularly, we focused our method on the important scenario of human figure splicing detection, and comprehensive experiments on multiple data sets and typical images proved the encouraging effectiveness of the new forensic clue and the proposed approach. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Improving spatial image adaptive steganalysis incorporating the embedding impactont he featureabstractRecently, in order to attack the adaptive steganograhpy more accurately, steganalysis features are associated with the content adaptivity. The adaptive σ version of the steganalysis features incorporates the impact of embedding on the residual to improve the detection. However, this method does not consider whether the embedding impact brings the change on the feature (histogram in the PSRM) which will be utilized by the detectors. Thus, we calculate the expectation of the residual L1distortion under the condition when the corresponding stego and cover residual values are within different quantization intervals, which will be accumulated in the histograms. This adaptive steganalytic scheme, with the relative position of the residual value in the quantization interval, only utilizes the residual distortion that leads to the change on the final feature. The experimental results demonstrate the potential of the proposed idea, especially for small payloads. This idea can also be applied to JPEG phase-aware features. Qingxiao Guan, Xianfeng Zhao, Jing Dong 0003, Zhoujun Xu |
ICIP | 4 |
| 2017 | Fragile image watermarking with pixel-wise recovery based on overlapping embedding strategy
Chuan Qin 0001, Ping Ji 0004, Xinpeng Zhang 0001, Jing Dong 0003 |
Signal Process. | 4 |
| 2017 | Optimized 3D Lighting Environment Estimation for Image Forgery DetectionabstractImage forgery is becoming a growing threat to information credibility. Among all kinds of image forgeries, photographic composites of human faces have very serious impacts. To combat this kind of forgery, some forensic methods propose to estimate the 3D lighting environments from different faces and investigate the consistency between them. Although they are very effective, existing 3D lighting-based forensic methods are limited by many simplifying assumptions about the surface reflection model, among which convexity and constant reflectance are two critical ones. In this paper, we propose an optimized 3D lighting estimation method by incorporating a more general surface reflection model. In this model, we relax the convexity and constant reflectance assumptions by taking the occlusion geometry and surface texture information into consideration. The proposed reflection model is more general and accurate; hence, it can achieve better lighting estimation accuracy and more reliable discrimination performance. Comprehensive experiments on both synthetic and real data sets validate the correctness and efficacy of the proposed method. Comparisons with two existing 3D lighting-based forensic methods also demonstrate the superiority of the proposed method for detecting face splicing. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Automatic detection of 3D lighting inconsistencies via a facial landmark based morphable modelabstractExisting 3D lighting consistency based forensic methods have some practical problems. They usually require additional images and human labor to reconstruct the 3D face model for lighting estimation, and furthermore, they cannot deal with expressional faces effectively. These drawbacks make them unusable in many practical cases. In this paper, we propose a more practical 3D lighting based forensic method by incorporating a facial landmark based 3D morphable model to efficiently fit the face shape. We also introduce a residual error based algorithm to automatically exclude outliers in lighting estimation. Our proposed method is fully automatic and very efficient compared to previous ones. Also, it does not depend on additional images and has better performance for expressional faces. Experiments on a realistic face dataset with variational lighting conditions indicate the efficacy and superiority of our method. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 3 |
| 2016 | Learning and transferring representations for image steganalysis using convolutional neural networkabstractThe major challenge of machine learning based image steganalysis lies in obtaining powerful feature representations. Recently, Qian et al. have shown that Convolutional Neural Network (CNN) is effective for learning features automatically for steganalysis. In this paper, we follow up this new paradigm in steganalysis, and propose a framework based on transfer learning to help the training of CNN for steganalysis, hence to achieve a better performance. We show that feature representations learned with a pre-trained CNN for detecting a steganographic algorithm with a high payload can be efficiently transferred to improve the learning of features for detecting the same steganographic algorithm with a low pay-load. By detecting representative WOW and S-UNIWARD steganographic algorithms, we demonstrate that the proposed scheme is effective in improving the feature learning in CNN models for steganalysis. Yinlong Qian, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 2 |
| 2015 | Robust steganalysis based on training set construction and ensemble classifiers weightingabstractThe cover source mismatch problem in steganalysis is a serious problem which keeps current steganalysis from practical use. It is mainly because of the high intra-class variation of cover and stego samples in the feature space, since current steganalytic features are inevitably affected much by the image content, size, quality and many other factors. Small training set often reflects only part of the real data distribution, hence the classifier (steganalyzer) may be undertrained and lack of robustness. In this paper, we propose a scheme to efficiently construct large representative training set for steganalysis. We also scheme out weighted ensemble classifiers which can be adaptive to testing data. Experimental results show that our method can improve the performance and robustness of ste-ganalysis under high intra-class variation. Xikai Xu, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 2 |
| 2014 | An effective watermarking method against valumetric distortionsabstractMost of the quantization based watermarking algorithms are very sensitive to valumetric distortions, while these distortions are regarded as common processing in audio/video analysis. In recent years, watermarking methods which can resist this kind of distortions have attracted a lot of interests. But still many proposed methods can only deal with one certain kind of valumetric distortion as amplitude scaling, and fail in other kinds of valumetric distortions like constant change attack or gamma correction. In this paper, we propose a very simple method to tackle all the three kinds of valumetric distortions. A constant change invariant domain is first constructed by spread transform, in which the watermark is embedded using a certain amplitude scaling invariant based watermarking scheme. Several typical watermarking methods and attacks have been implemented in our experiments to demonstrate the effectiveness of the proposed method. Zairan Wang, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 2 |
| 2014 | Effects of Fragile and Semi-fragile Watermarking on Iris Recognition System
Zairan Wang, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
IWDW | 2 |
| 2014 | Exploring DCT Coefficient Quantization Effects for Local Tampering DetectionabstractIn this paper, we focus on local image tampering detection. For a JPEG image, the probability distributions of its DCT coefficients will be disturbed by tampering operation. The tampered region and the unchanged region have different distributions, which is an important clue for locating tampering. Based on the assumption of Laplacian distribution of unquantized ac DCT coefficients, these two distributions as well as the size of tampered region can be estimated so that the probability of each DCT block being tampered is obtained. More accurate localization results could be got when we consider the prior knowledge of common tampered regions. We also design three kinds of features that can distinguish truly tampered regions from the false ones to reduce false alarm. For a tampered image which is saved in lossless compressed format, we also propose the specialized approach, which employs the quantization noise of high-frequency DCT coefficient, to improve the tampering localization performance. Extensive experiments on large scale databases prove the effectiveness of our proposed method and demonstrate that our method is suitable for locating tampered regions with different scales. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Two Notes from Experimental Study on Image Steganalysis
Qingxiao Guan, Jing Dong 0003, Tieniu Tan |
ICIC (1) | 2 |
| 2013 | Video steganalysis based on the constraints of motion vectorsabstractIn this paper, we focus on detecting data hiding in motion vectors of compressed video and propose a new steganalytic algorithm based on the mutual constraints of motion vectors. The constraints of motion vectors from multiple frames are analyzed and formulized by three functions, then statistical features are extracted based on these functions. Moreover, we also incorporate calibration method to improve the detection accuracy. Experimental results demonstrate that the proposed method can effectively attack typical motion-vector-based video steganography. Xikai Xu, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 2 |
| 2012 | Universal spatial feature set for video steganalysisabstractIn this paper, we propose a universal spatial feature set for video steganalysis. This feature set comprehensively exploits the correlation of adjacent pixels and can be viewed as a generalized extension of most correlation based features. We also develop a new approach to extract inter-frame features for video steganalysis. Our method can be universally applied to detect different video steganographic algorithms regardless of video format. The experimental results show that it outperforms current correlation based methods. Xikai Xu, Jing Dong 0003, Tieniu Tan |
ICIP | 2 |
| 2011 | An effective image steganalysis method based on neighborhood information of pixelsabstractThis paper focuses on image steganalysis. We use higher order image statistics based on neighborhood information of pixels (NIP) to detect the stego images from original ones. We use subtracting gray values of adjacent pixels to capture neighborhood information, and also make use of “rotation invariant” property to reduce the dimensionality for the whole feature sets. We tested two kinds of NIP feature, the experimental results illustrates that our proposed feature sets are with good performance and even outperform the state-of-art in certain aspect. Qingxiao Guan, Jing Dong 0003, Tieniu Tan |
ICIP | 2 |
| 2010 | Image tampering detection based on stationary distribution of Markov chainabstractIn this paper, we propose a passive image tampering detection method based on modeling edge information. We model the edge image of image chroma component as a finite-state Markov chain and extract low dimensional feature vector from its stationary distribution for tampering detection. The support vector machine (SVM) is utilized as classifier to evaluate the effectiveness of the proposed algorithm. The experimental results in a large scale of evaluation database illustrates that our proposed method is promising. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 2 |
| 2010 | New developments in color image tampering detectionabstractIn this paper, an efficient framework for passive-blind color image tampering detection is presented. Statistical features are extracted from a given test image and a set of 2-D arrays derived by applying multi-size block discrete cosine transform to the given test image. Image features are extracted from Cr channel, a chroma channel in YCbCr color space, because of its observed sensitivity to color image tampering. A support vector machine is employed to evaluate the effectiveness of image features over a color image dataset recently established for tampering detection. Boosting feature selection is applied to having feature dimensionality reduced so as to make detection accuracy generalizable and computational complexity decreased. Experimental results have demonstrated that the proposed framework applied to the aforementioned dataset outperforms the state of the arts by distinct margins. Patchara Sutthiwan, Yun Q. Shi 0001, Jing Dong 0003, Tieniu Tan, Tian-Tsong Ng |
ISCAS | 3 |
| 2010 | Blind Quantitative Steganalysis Based on Feature Fusion and Gradient Boosting
Qingxiao Guan, Jing Dong 0003, Tieniu Tan |
IWDW | 2 |
| 2010 | Tampered Region Localization of Digital Color Images Based on JPEG Compression Noise
Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IWDW | 2 |
| 2009 | Effective image splicing detection based on image chromaabstractA color image splicing detection method based on gray level co-occurrence matrix (GLCM) of thresholded edge image of image chroma is proposed in this paper. Edge images are generated by subtracting horizontal, vertical, main and minor diagonal pixel values from current pixel values respectively and then thresholded with a predefined threshold T. The GLCMs of edge images along the four directions serve as features for image splicing detection. Boosting feature selection is applied to select optimal features and Support Vector Machine (SVM) is utilized as classifier in our approach. The effectiveness of the proposed method has been demonstrated by our experimental results. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 2 |
| 2009 | Multi-class Blind Steganalysis Based on Image Run-Length Analysis
Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
IWDW | 1 |
| 2009 | A Survey of Passive Image Tampering Detection
Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IWDW | 2 |
| 2008 | Effects of watermarking on iris recognition performanceabstractProtection of biometric data and templates is a crucial issue for the security of biometric systems, and biometric watermarking is introduced for this purpose. However, watermarking introduces extra information into the biometric data (biometric images or biometric feature templates) which leads to certain distortion. In addition, watermarked images are always subject to the risk of being attacked. Hence, whether and how biometric recognition performance will be affected by biometric watermarking deserves investigation. In this paper, we make a first attempt in such investigations by studying two application scenarios in the context of iris recognition, namely protection of iris templates by hiding them in cover images as watermarks (iris watermarks), and protection of iris images by watermarking them. Experimental results suggest that watermark embedding in iris images does not introduce detectable decreases on iris recognition performance whereas recognition performance drops significantly if iris watermarks suffer from severe attacks. Jing Dong 0003, Tieniu Tan |
ICARCV | 1 |
| 2008 | Blind image steganalysis based on run-length histogram analysisabstractIn this paper, a new, simple but effective method is proposed for blind image steganalysis, which is based on run-length histogram analysis. Higher-order statistics of characteristic functions of three types of image run-length histograms are selected as features. Support vector machine is used as classifier. Experimental results demonstrate that the proposed scheme significantly outperforms prior arts in detection accuracy and generality. Jing Dong 0003, Tieniu Tan |
ICIP | 1 |
| 2008 | Run-Length and Edge Statistics Based Approach for Image Splicing Detection
Jing Dong 0003, Wei Wang 0025, Tieniu Tan, Yun Q. Shi 0001 |
IWDW | 1 |
| 2007 | Fusion Based Blind Image Steganalysis by Boosting Feature Selection
Jing Dong 0003, Xiaochuan Chen, Tieniu Tan |
IWDW | 1 |