VLDB 2026 Research / reviewers in the wild / expert
Jian Zhang 0086
dblp:07/314-86
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0009-0007-9632-3785ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | F2Mamba: Forgery-Guided Vision Mamba With Multi-Scale Frequency Perception for General Image Forgery Localization
Jiangqun Ni, Fan Nie, Jian Zhang 0086 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Toward Generalizable Deepfake Detection via Forgery-Aware Audio-Visual Adaptation: A Variational Bayesian ApproachabstractThe widespread application of AIGC contents has brought not only unprecedented opportunities, but also potential security concerns, e.g., audio-visual deepfakes. Therefore, it is of great importance to develop an effective and generalizable method for multi-modal deepfake detection. Typically, the audio-visual correlation learning could expose subtle cross-modal inconsistencies, e.g., audio-visual misalignment, which serve as crucial clues in deepfake detection. In this paper, we reformulate the correlation learning with variational Bayesian estimation, where audio-visual correlation is approximated as a Gaussian distributed latent variable, and thus develop a novel framework for deepfake detection, i.e., Forgery-aware Audio-Visual Adaptation with Variational Bayes (FoVB). Specifically, given the prior knowledge of pre-trained backbones, we adopt two core designs to estimate audio-visual correlations effectively. First, we exploit various difference convolutions and a high-pass filter to discern local and global forgery traces from both modalities. Second, with the extracted forgery-aware features, we estimate the latent Gaussian variable of audio-visual correlation via variational Bayes. Then, we factorize the variable into modality-specific and correlation-specific ones with orthogonality constraint, allowing them to better learn intra-modal and cross-modal forgery traces with less entanglement. Extensive experiments demonstrate that our FoVB outperforms other state-of-the-art methods in various benchmarks. Fan Nie, Jiangqun Ni, Jian Zhang 0086, Bin Zhang 0048, Weizhe Zhang, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Forgery-Aware and Edge-Guided Diffusion Model for General and Robust Image Forgery LocalizationabstractImage forgery localization, which aims to find suspicious tampered regions by splicing, copy-move, or removal manipulations, has attracted increasing attention. Although considerable progress has been made, most of the existing methods are still far from satisfactory in terms of generalization, e.g., cross-dataset evaluation, and show less robustness against lossy distortion, e.g., post-processing attacks or transmission over online social networks. To address these challenges, a general and robust image forgery localization framework using a diffusion probabilistic model is proposed in this paper. First, aDiffusion-drivenForgery-awareLocalization network (DFL) is presented. In specific, the task of image forgery localization is formulated as the one of mask reconstruction with the forgery-aware features as conditional prior, which introduces forgery-related knowledge into the diffusion process to gradually recover the ground-truth mask from the noisy one. AnEdge-guidedDiffusionRestoration network (EDR) is then developed and integrated into the DFL network, leading to theEDR-DFLmodel, boosting the performance against various post-processing attacks. With EDR-DFL, the distorted forged images could be restored using the EDR model in an edge feature preserving way, which allows the DFL model to localize the tampered regions in the restored images effectively. Extensive experimental results demonstrate that the proposed method significantly outperforms other state-of-the-art methods for cross-dataset evaluation and exhibits superior robustness against various challenging post-processing attacks. Jiangqun Ni, Jian Zhang 0086 |
IEEE Trans. Multim. | 3 |
| 2025 | Reinforced Multi-teacher Knowledge Distillation for Efficient General Image Forgery Detection and LocalizationabstractImage forgery detection and localization (IFDL) is of vital importance as forged images can spread misinformation that poses potential threats to our daily life. However, previous methods still struggled to effectively handle forged images processed with diverse forgery operations in real-world scenarios. In this paper, we propose a novel Reinforced Multi-teacher Knowledge Distillation (Re-MTKD) framework for the IFDL task, structured around an encoder-decoder ConvNeXt-UperNet along with Edge-Aware Module, named Cue-Net. First, three Cue-Net models are separately trained for the three main types of image forgeries, i.e., copy-move, splicing and inpainting, which then serve as the multi-teacher models to train the target student model with Cue-Net through self-knowledge distillation. A Reinforced Dynamic Teacher Selection (Re-DTS) strategy is developed to dynamically assign weights to the involved teacher models, which facilitates specific knowledge transfer and enables the student model to effectively learn both the common and specific natures of diverse tampering traces. Extensive experiments demonstrate that, compared with other state-of-the-art methods, the proposed method achieves superior performance on several recently emerged datasets comprised of various kinds of image forgeries. Zeqin Yu, Jiangqun Ni, Jian Zhang 0086, Haoyi Deng, Yuzhen Lin |
AAAI | 3 |
| 2025 | JPEG-RAE: Reversible Adversarial Example for Privacy and Copyright Protection of JPEG ImagesabstractReversible Adversarial Example (RAE) could be used to protect the privacy and copyright of images on social networks (SONs) by exploring the adversarial examples to disrupt the access of malicious AI models while ensuring recoverability with authorized users. Existing RAE methods add adversarial perturbations in spatial images which do not apply to JPEG images, the most widely adopted image format for image storage and transmission. To tackle this issue, we propose the first Reversible Adversarial Example (JPEG-RAE) generation framework for JPEG images, which consists of two primary components, i.e., JPEG-AE and G-RDH. JPEG-AE crafts the adversarial perturbations in the JPEG domain of images by leveraging chain rule of gradient propagation, so that they could effectively mislead the AI models in spatial domain when they are JPEG decompressed. And G-RDH adopts a gradient-directed bi-directional histogram shifting scheme for efficient reversible hiding of adversarial perturbations and location data in JPEG domain, where the histogram shifting is in sync with the sign of back-propagated gradients to further boost the performance of adversarial attacks. Experimental validation demonstrates that, although confined to the JPEG format such as the amount and intensity of alterable DCT coefficients, the proposed JPEG-RAE could still show superior or comparable performance, in terms of attack ability and recover ability, to its counterparts in spatial domain. Dahao Fu, Jiangqun Ni, Jian Zhang 0086 |
ACM Multimedia | 3 |
| 2025 | DSM: Domain Shift Modeling for general deepfake detection
Jian Zhang 0086, Jiangqun Ni, Fan Nie |
Signal Process. | 1 |
| 2025 | DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake DetectionabstractWith the advancement of deepfake generation techniques, the importance of deepfake detection in protecting multimedia content integrity has become increasingly obvious. Recently, temporal inconsistency clues have been explored to improve the generalizability of deepfake video detection. According to our observation, the temporal artifacts of forged videos in terms of motion information usually exhibits quite distinct inconsistency patterns along horizontal and vertical directions, which could be leveraged to improve the generalizability of detectors. In this paper, a transformer-based framework forDiffusion Learning ofInconsistencyPattern (DIP) is proposed, which exploits directional inconsistencies for deepfake video detection. Specifically, DIP begins with a spatiotemporal encoder to represent spatiotemporal information. A directional inconsistency decoder is adopted accordingly, where direction-aware attention and inconsistency diffusion are incorporated to explore potential inconsistency patterns and jointly learn the inherent relationships. In addition, the SpatioTemporal Invariant Loss (STI Loss) is introduced to contrast spatiotemporally augmented sample pairs and prevent the model from overfitting nonessential forgery artifacts. Extensive experiments on several public datasets demonstrate that our method could effectively identify directional forgery clues and achieve state-of-the-art performance. Fan Nie, Jiangqun Ni, Jian Zhang 0086, Bin Zhang 0048, Weizhe Zhang |
IEEE Trans. Multim. | 3 |
| 2025 | Domain-invariant and Patch-discriminative Feature Learning for General Deepfake DetectionabstractHyper-realistic avatars in the metaverse have already raised security concerns about deepfake techniques; deepfakes involving generated video “recording” may be mistaken for a real recording of the people it depicts. As a result, deepfake detection has drawn considerable attention in the multimedia forensic community. Though existing methods for deepfake detection achieve fairly good performance under the intra-dataset scenario, many of them gain unsatisfying results in the case of cross-dataset testing with more practical value, where the forged faces in training and testing datasets are from different domains. To tackle this issue, in this article, we propose a novel Domain-Invariant and Patch-Discriminative feature learning framework—DI&PD. For image-level feature learning, a single-side adversarial domain generalization is introduced to eliminate domain variances and learn domain-invariant features in training samples from different manipulation methods, along with the global and local random crop augmentation strategy to generate more data views of forged images at various scales. A graph structure is then built by splitting the learned image-level feature maps, with each spatial location corresponding to a local patch, which facilitates patch representation learning by message-passing among similar nodes. Two types of center losses are utilized to learn more discriminative features in both image-level and patch-level embedding spaces. Extensive experimental results on several datasets demonstrate the effectiveness and generalization of the proposed method compared with other state-of-the-art methods. Jian Zhang 0086, Jiangqun Ni, Fan Nie, Jiwu Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Diff-IFL: Towards General Image Forgery Localization using Diffusion Probabilistic ModelabstractWith wide applications of image editing tools, forged images have become a great public concern. Although existing methods for image forgery localization (IFL) could achieve fairly good results on several public datasets, most of them perform unsatisfactorily for cross-dataset evaluation and online social network applications. To tackle this issue, a novel coarse-to-fine framework using Diffusion probabilistic model for Image Forgery Localization (Diff-IFL) is proposed in this paper, which consists of a coarse localization module and a mask diffusion module. The coarse localization module employs a transformer-based architecture to represent the tampered images and generate coarse masks. While the mask diffusion module formulates IFL as a mask reconstruction task, it relies on the extracted forgery feature representations as the conditional prior to gradually recover the clean ground-truth mask from the noisy mask. Extensive experiments demonstrate that Diff-IFL outperforms other SOTA methods and exhibits superior robustness against social media forgery. Jiangqun Ni, Jian Zhang 0086, Shiyuan Tang |
ICME | 3 |
| 2024 | FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake DetectionabstractNowadays, the abuse of AI-generated content (AIGC), especially the facial images known as deepfake, on social networks has raised severe security concerns, which might involve the manipulations of both visual and audio signals. For multimodal deepfake detection, previous methods usually exploit forgery-relevant knowledge to fully finetune Vision transformers (ViTs) and perform cross-modal interaction to expose the audio-visual inconsistencies. However, these approaches may undermine the prior knowledge of pretrained ViTs and ignore the domain gap between different modalities, resulting in unsatisfactory performance. To tackle these challenges, in this paper, we propose a new framework, i.e., Forgery-aware Audio-distilled Multimodal Learning (FRADE), for deepfake detection. In FRADE, the parameters of pretrained ViT are frozen to preserve its prior knowledge, while two well-devised learnable components, i.e., the Adaptive Forgery-aware Injection (AFI) and Audio-distilled Cross-modal Interaction (ACI), are leveraged to adapt forgery relevant knowledge. Specifically, AFI captures high-frequency discriminative features on both audio and visual signals and injects them into ViT via the self-attention layer. Meanwhile, ACI employs a set of latent tokens to distill audio information, which could bridge the domain gap between audio and visual modalities. The ACI is then used to well learn the inherent audio-visual relationships by cross-modal interaction. Extensive experiments demonstrate that the proposed framework could outperform other state-of-the-art multimodal deepfake detection methods under various circumstances. Fan Nie, Jiangqun Ni, Jian Zhang 0086, Bin Zhang 0048, Weizhe Zhang |
ACM Multimedia | 3 |
| 2024 | EFLNet: Enhancing Feature Learning Network for Infrared Small Target DetectionabstractSingle-frame infrared small target detection is considered to be a challenging task, due to the extreme imbalance between target and background, bounding box regression is extremely sensitive to infrared small target, and target information is easy to lose in the high-level semantic layer. In this article, we propose an enhancing feature learning network (EFLNet) to address these problems. First, we notice that there is an extremely imbalance between the target and the background in the infrared image, which makes the model pay more attention to the background features rather than target features. To address this problem, we propose a new adaptive threshold focal loss (ATFL) function that decouples the target and the background, and utilizes the adaptive mechanism to adjust the loss weight to force the model to allocate more attention to target features. Second, we introduce the normalized Gaussian Wasserstein distance (NWD) to alleviate the difficulty of convergence caused by the extreme sensitivity of the bounding box regression to infrared small target. Finally, we incorporate a dynamic head mechanism into the network to enable adaptive learning of the relative importance of each semantic layer. Experimental results demonstrate our method can achieve better performance in the detection performance of infrared small target compared to the state-of-the-art (SOTA) deep-learning-based methods. The source codes and bounding box annotated datasets are available athttps://github.com/YangBo0411/infrared-small-target. Jian Zhang 0086, Jun Luo 0006, Mingliang Zhou 0001, Yangjun Pi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Domain-Invariant Feature Learning for General Face Forgery DetectionabstractThough existing methods for face forgery detection achieve fairly good performance under the intra-dataset scenario, few of them gain satisfying results in the case of cross-dataset testing with more practical value. To tackle this issue, in this paper, we propose a novel domain-invariant feature learning framework - DIFL for face forgery detection. In the framework, an adversarial domain generalization is introduced to learn the domain-invariant features from the forged samples synthesized by various algorithms. Then a center loss in fractional form (CL) is utilized to learn more discriminative features by aggregating the real faces while separating the fake faces from the real ones in the embedding space. In addition, a global and local random crop augmentation strategy is utilized to generate more data views of forged facial images at various scales. Extensive experimental results demonstrate the effectiveness and generalization of the proposed method compared with other state-of-the-art methods. Jian Zhang 0086, Jiangqun Ni |
ICME | 1 |
| 2022 | Evading generated-image detectors: A deep dithering approach
Hao Xie 0002, Jiangqun Ni, Jian Zhang 0086, Weizhe Zhang, Jiwu Huang |
Signal Process. | 3 |
| 2022 | Corrigendum to 'Evading generated-image detectors: A deep dithering approach' [Signal Processing 197(2022) 108558]
Hao Xie 0002, Jiangqun Ni, Jian Zhang 0086, Weizhe Zhang, Jiwu Huang |
Signal Process. | 3 |
| 2021 | DeepFake Videos Detection Using Self-Supervised Decoupling NetworkabstractWith wide applications of facial manipulation technology, fake images and videos are becoming a great public concern. Although existing methods for face forgery detection could achieve fairly good results on public database, most of them perform poorly when the fake images/videos are compressed as they are usually done in social networks. To tackle this issue, a self-supervised decoupling network (SSDN), that incorporates compression irrelevance, is proposed in this paper. The proposed model learns two separate feature representations for the suspect videos, i.e., authenticity and compression. A joint self-supervised strategy is then utilized for feature decoupling, in which, the similarity decoupling is carried out by similarity learning on authentic features, whereas for adversarial decoupling, the proposed SSDN model is trained in an adversarial manner for robust feature learning. Experimental results show that the SSDN outperforms the state-of-the-art methods for deepfake detection against compression attacks on public datasets, e.g., FaceForensics++. Jian Zhang 0086, Jiangqun Ni, Hao Xie 0002 |
ICME | 1 |