EDBT 2026 Demo / reviewers in the wild / expert
Jiacheng Deng 0001
dblp:320/4938-1
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-0452-4031ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time Shuffle: A Transferability-Booster for Multiple Audio Adversarial TasksabstractExisting audio adversarial attack methods suffer from poor transferability, primarily due to insufficient exploration of model decision mechanisms and overreliance on heuristic-driven algorithm design. This paper aims to alleviate this gap. Specifically, through observations across three mainstream audio tasks (Automatic Speech Recognition, Speaker Verification, and Keyword Spotting), we reveal that these models primarily rely on local temporal features—inputs with time shuffled retain 83.7% of original accuracy. The SHAP-based visualization further validated that time shuffle leads to a significant shift in the salient regions of the model, but the samples can still be correctly identified, indicating the presence of redundant features that can affect decision-making. Inspired by these findings, we propose Time-Shuffle (TS) adversarial attack (including segments-based TS and phoneme-level-based TS-p). This method divides audio or phonemes into segments, randomly shuffles them, and computes gradients on the shuffled structure. By forcing perturbations to exploit transferable local temporal features and reduce overfitting to source-specific patterns, TS/TS-p inherently enhances transferability. As a model-agnostic framework, TS/TS-p can seamlessly integrate with existing attack methods. Comprehensive experiments demonstrate that TS-p achieved SOTA and boosts transferability by about 23%/14.7%/6.3% on ASR/ASV/KWS. Jiacheng Deng 0001, Dengpan Ye, Zhaolin Wei, Ziyi Liu 0009 |
AAAI | 1 |
| 2026 | Perceptual V-Cloak: Generating Trainable and Unintelligible Speech Dataset
Dengpan Ye, Jiacheng Deng 0001, Zhaolin Wei, Ziyi Liu 0009 |
ICIC (2) | 3 |
| 2026 | MSFT-Net: Mixture Semantic-Agnostic Manipulation Trace Enhanced Architecture for Robust Image Manipulation LocalizationabstractSince the proliferation of image manipulation methods, effective image manipulation localization (IML) in scenarios with post-processing operations gradually becomes a core challenge. For a long time, IML either relies on strongly semantically related features, resulting in semantic relevance bias in the localization results, or only uses a single semantic-agnostic space feature, which is unable to maintain effective localization capabilities after image post-processing operations. Inspired by this, we propose a novel mixture semantic-agnostic manipulation trace robust localization network (MSFT-Net), which specifically utilizes mixture semantic-agnostic information to achieve effective and robust IML. The MSFT-Net introduces two new modules, the mixture shared manipulation trace enhancement module (MISE) and the Multiscale Feature Association Module (FAM). MISE dynamically links multiple semantic-agnostic feature extractors using a sparsity-enhanced mixture of shared experts, enabling the extraction of diverse manipulation features for accurate localization. Furthermore, keeping the high resolution of the localization features is very important in the mask prediction stage. Therefore, FAM outputs high-resolution fused manipulation features by using the correlation of features at the same level and the spatial context information from different levels. This further improves the effectiveness of IML in post-processing scenarios. Comprehensive experiments on five datasets demonstrate that our model significantly improves both in localization accuracy (average F1 score and IoU increasing by over 9.9% and 4.0%) and robustness. The codes will be made available. Dengpan Ye, Yunming Zhang, Jiacheng Deng 0001, Ziyi Liu 0009, Yueyun Shang, Zhihong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Take Fake as Real: Realistic-Like Robust Black-Box Adversarial Attack to Evade AIGC DetectionabstractThe security of AI-generated content (AIGC) detection is crucial for ensuring multimedia content credibility. To enhance detector security, research on adversarial attacks has become essential. However, most existing adversarial attacks focus only on GAN-generated facial images detection, struggle to be effective on multi-class natural images and diffusion-based detectors, and exhibit poor invisibility. To fill this gap, we first conduct an in-depth analysis of the vulnerability of AIGC detectors and discover the feature that detectors vary in vulnerability to different post-processing. Then, considering that the detector is agnostic in real-world scenarios and given this discovery, we propose a Realistic-like Robust Black-box Adversarial attack (R2BA) with post-processing fusion optimization. Unlike typical perturbations, R2BA uses real-world post-processing, i.e., Gaussian blur, JPEG compression, Gaussian noise and light spot to generate adversarial examples. Specifically, we use a stochastic particle swarm algorithm with inertia decay to optimize post-processing fusion intensity and explore the detector’s decision boundary. Guided by the detector’s fake probability, R2BA enhances/weakens the detector-vulnerable/detector-robust post-processing intensity to strike a balance between adversariality and invisibility. Extensive experiments on popular/commercial AIGC detectors and datasets demonstrate that R2BA exhibits impressive anti-detection performance, excellent invisibility, and strong robustness in GAN-based and diffusion-based cases. Compared to state-of-the-art white-box and black-box attacks, R2BA shows significant improvements of 15%–72% and 21%–47% in anti-detection performance under the original and robust scenario respectively, offering valuable insights for the security of AIGC detection in real-world applications. Caiyun Xie, Dengpan Ye, Yunming Zhang, Yueyun Shang, Yunna Lv, Jiacheng Deng 0001, Jiawei Song |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | DIP-Watermark: A Double Identity Protection Method Based on Robust Adversarial WatermarkabstractThe wide deployment of Face Recognition (FR) systems poses privacy risks. One countermeasure is adversarial attack, deceiving unauthorized malicious FR, but it also disrupts regular identity verification of trusted authorizers, exacerbating the potential threat of identity impersonation. To address this, we propose the first double identity protection scheme based on traceable adversarial watermarking, termed DIP-Watermark. DIP-Watermark employs a one-time watermark embedding to deceive unauthorized FR models and allows authorizers to perform identity verification by extracting the watermark. Specifically, we propose an information-guided adversarial attack against FR models. The encoder embeds an identity-specific watermark into the deep feature space of the carrier, guiding recognizable features of the image to deviate from the source identity. We further adopt a collaborative meta-optimization strategy compatible with sub-tasks, which regularizes the joint optimization direction of the encoder and decoder. This strategy enhances the representation of universal carrier features, mitigating multi-objective optimization conflicts in watermarking. Extensive experiments on two large-scale facial datasets demonstrate that DIP-Watermark achieves significant attack success rates and traceability accuracy on state-of-the-art FR models and commercial APIs. It also exhibits superior robustness against a wide range of real-world simulated distortions, outperforming existing privacy protection methods based on adversarial attacks, deep watermarking, or their simple combination. Our work potentially opens up new insights into proactive protection for FR privacy. Yunming Zhang, Dengpan Ye, Caiyun Xie, Sipeng Shen, Ziyi Liu 0009, Jiacheng Deng 0001, Yueyun Shang, Zhihong Tian 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Generalize Audio Deepfake Algorithm Recognition via Attribution EnhancementabstractThe development of voice cloning techniques has made forgery audios indistinguishable, posing an urgency to trace their sources. Many existing works focus on improving identification accuracy for audio deepfake algorithm recognition. However, most methods ignore the impact of complex information in audio signals on attribution. In this paper, we propose an audio deepfake attribution enhancement (ADAE) strategy, which aims to magnify the fingerprints of generation styles by removing the speaker information. This is achieved through an information disentangle block with an extra speaker encoder. In addition, we propose the FakeSource dataset, a novel audio deepfake algorithm recognition dataset that contains 25 different voice cloning algorithms, to address the constraint of data scarcity. Experiments on the FakeSource dataset demonstrate that ADAE improves the performance of unseen algorithm detection. We also assess ADAE on a more challenging training-free task which shows competitive performance. Dengpan Ye, Jiacheng Deng 0001 |
ICASSP | 4 |
| 2025 | From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in BackgroundabstractMusic deepfake detection is aimed at identifying whether songs are generated by AI. Current methods usually separate vocals from background music for detection, but this could leave residual forgery information in the background. Our study demonstrates for the first time that incorporating background forgery information with vocals can improve detection accuracy. Furthermore, our findings show that using background features can reduce EER by an average of about 2% on existing frameworks. Based on this observation, we propose a novel Hybrid Frontend that captures generalized features from both vocal and background music. The Hybrid Frontend comprises two branches: vocal and background music part. Specifically, the vocal part uses the Sinconv encoder as a deeply embedded feature extractor. The latter captures background variation by fine-tuning the pre-trained model with adapters. Experimental results demonstrate that our method outperforms the vocal-only detection on WildSVDD dataset, achieving an EER of 8.53%, which is 1.3% lower. Zhaolin Wei, Dengpan Ye, Jiacheng Deng 0001 |
ICASSP | 3 |
| 2025 | PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial Perturbations
Zhaolin Wei, Xiuwen Shi, Dengpan Ye, Jiacheng Deng 0001, Ziyi Liu 0009 |
ACM Multimedia | 6 |
| 2025 | Three-in-One: Robust Enhanced Universal Transferable Anti-Facial Retrieval in Online Social NetworksabstractDeep hash-based retrieval techniques are widely used in facial retrieval systems to improve the efficiency of facial matching. However, it also carries the danger of exposing private information. Deep hash models are easily influenced by adversarial examples, which can be leveraged to protect private images from malicious retrieval. The existing adversarial example methods against deep hash models focus on universality and transferability, lacking the research on its robustness in online social networks (OSNs), which leads to their failure in anti-retrieval after post-processing. Therefore, we provide the first in-depth discussion on robustness in universal transferable anti-facial retrieval and propose Three-in-One Adversarial Perturbation (TOAP). Specifically, we construct a local and global Compression Generator (CG) to simulate complex post-processing scenarios, which can be used to mitigate perturbation. Then, we propose robust optimization objectives based on the discovery of the variation patterns of model’s distribution after post-processing, and generate adversarial examples using these objectives and meta-learning. Finally, we iteratively optimize perturbation by alternately generating adversarial examples and fine-tuning the CG, balancing the performance of perturbation while enhancing CG’s ability to mitigate them. Numerous experiments demonstrate that, in addition to its advantages in universality and transferability, TOAP significantly outperforms current state-of-the-art methods in multiple robustness metrics. It further improves universality and transferability by 5% to 28%, and achieves up to about 33% significant improvement in several simulated post-processing scenarios as well as mainstream OSNs, demonstrating that TOAP can effectively protect private images from malicious retrieval in real-world scenarios. Yunna Lv, Dengpan Ye, Caiyun Xie, Jiacheng Deng 0001, Yiheng He, Sipeng Shen |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | The Interpretable and Transferable Adversarial Attack against Synthetic Speech DetectorsabstractExisting work finds it challenging for adversarial examples to transfer among different synthetic speech detectors because of cross-feature and cross-model. To enhance the transferability of adversarial examples, we propose a spectral saliency analysis method and gain insight into the underlying detection mechanisms of existing detectors for the first time. These insights offer an interpretable basis for why adversarial examples are challenging to transfer between synthetic speech detection models. Then we further propose a two-stage adversarial attack framework. Specifically, the first stage leverages insights into the model detection mechanism to design a random time-frequency masking module, the random offset module, and 1D convolution to generate transferable and robust adversarial examples. In the second stage, to mitigate the problem of obvious noise in the low-energy frames of the carrier in existing adversarial attacks, we perform secondary optimization on frames below the Signal-Noise-Rate threshold to enhance its auditory quality. Extensive experimental results demonstrate that the proposed method significantly enhances the transferability and robustness of adversarial examples, while simultaneously preserving the acoustic quality compared to typical approaches. Jiacheng Deng 0001, Dengpan Ye, Jizhi Li, Ziyi Liu 0009, Yunming Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | AVT$^{2}$-DWF: Improving Deepfake Detection With Audio-Visual Fusion and Dynamic Weighting StrategiesabstractWith the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this letter, we proposeAVT$^{2}$-DWF, theAudio-Visual dualTransformers grounded inDynamicWeightFusion, which aims to amplify both intra- and cross-modal forgery cues, thereby enhancing detection capabilities. AVT$^{2}$-DWF adopts a dual-stage approach to capture both spatial characteristics and temporal dynamics of facial expressions. This is achieved through a face transformer with an$n$-frame-wise tokenization strategy encoder and an audio transformer encoder. Subsequently, it uses multi-modal conversion with dynamic weight fusion to address the challenge of heterogeneous information fusion between audio and visual modalities. Experiments on DeepfakeTIMIT, FakeAVCeleb, and DFDC datasets indicate that AVT$^{2}$-DWF achieves state-of-the-art performance intra- and cross-dataset Deepfake detection. Rui Wang 0141, Dengpan Ye, Yunming Zhang, Jiacheng Deng 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Dual Defense: Adversarial, Traceable, and Invisible Robust Watermarking Against Face SwappingabstractMalicious applications of deep face swapping technology pose security threats such as misinformation dissemination and identity fraud. Some research propose the utilization of robust watermarking methods to track the copyright of facial images, facilitating post-forgery identity attribution. However, these methods cannot fundamentally prevent or eliminate the adverse impacts of face swapping. To address this issue, we present Dual Defense, an innovative framework based on robust adversarial watermarking. It simultaneously tracks image copyrights and disrupts the face swapping model by one-time embedding the robust adversarial watermark. Specifically, we propose an Original-domain Feature Emulation Attack (OFEA) method, which makes the traceable watermark adversarial through specially designed original domain adversarial loss. Additionally, we conduct a wavelet domain image structural information compensation loss, combined with a channel attention mechanism, to jointly balance watermark invisibility, adversariality, and traceability. Furthermore, we design a more comprehensive and rational evaluation method to thoroughly assess the effectiveness of adversarial attacks against face swapping models. Extensive experiments demonstrate that Dual Defense exhibits exceptional cross-task generality and dataset generalization. It maintains impressive adversariality and traceability in both original and robust settings, surpassing current forgery defense methods that possess only one of these capabilities. Yunming Zhang, Dengpan Ye, Caiyun Xie, Xin Liao 0001, Ziyi Liu 0009, Chuanxi Chen, Jiacheng Deng 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2023 | Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character RecognitionabstractOptical Character Recognition (OCR) enables automatic text extraction from scanned or digitized text images, but it also makes it easy to pirate valuable or sensitive text from these images. Previous methods to prevent OCR piracy by distorting characters in text images are impractical in real-world scenarios, as pirates can capture arbitrary portions of the text images, rendering the defenses ineffective. In this work, we propose a novel and effective defense mechanism termed the Universal Defensive Underpainting Patch (UDUP) that modifies the underpainting of text images instead of the characters. UDUP is created through an iterative optimization process to craft a small, fixed-size defensive patch that can generate non-overlapping underpainting for text images of any size. Experimental results show that UDUP effectively defends against unauthorized OCR under the setting of any screenshot range or complex image background. It is agnostic to the content, size, colors, and languages of characters, and is robust to typical image operations such as scaling and compressing. In addition, the transferability of UDUP is demonstrated by evading several off-the-shelf OCRs. The code is available at https://github.com/QRICKDD/UDUP. Jiacheng Deng 0001, Li Dong 0006, Diqun Yan, Rangding Wang, Dengpan Ye, Lingchen Zhao, Jinyu Tian 0001 |
ACM Multimedia | 1 |
| 2023 | Pseudo-label Diversity Exploitation for Few-Shot Object Detection
Chong Wang 0001, Zhengjie Ye, Jiacheng Deng 0001 |
MMM (2) | 5 |
| 2023 | Stealthy Backdoor Attack Against Speaker Recognition Using Phase-Injection Hidden TriggerabstractDeep learning has achieved significant breakthroughs in speaker recognition, driven by continual advancements in foundation models. However, malicious third-party platforms have introduced a severe security concern through backdoor attacks, in which attackers can manipulate a model to output a specific label by implanting a trigger. Existing speech backdoor attack methods typically utilize fixed and unnoticeable perturbations as triggers, but these may still be audible and thus detected during training and inference stages. To overcome this limitation, we propose a novel backdoor attack paradigm (PhaseBack) injecting triggers in the phase spectrum. PhaseBack exhibits sufficient stealth by leveraging the fact that the human ear is insensitive to phase information. Besides, injecting partial perturbations in the frequency domain results in global perturbations throughout the time domain, making the attack more effective. Extensive experiments on the Voxceleb1 dataset demonstrate the effectiveness and stealthiness of PhaseBack. Moreover, it has strong resistance to bypass several defense methods. Zhe Ye 0001, Diqun Yan, Li Dong 0006, Jiacheng Deng 0001, Shui Yu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Decision-Based Attack to Speaker Recognition System via Local Low-Frequency PerturbationabstractDespite neural network-based speaker recognition systems (SRS) have enjoyed significant success, they are proved to be quite vulnerable to adversarial examples. In practice, the SRS model parameters are not always available. Attackers have to probe the model only via querying, and such decision-based attacking merely relies on the output label is quite challenging. This letter proposes a two-step query-efficient decision-based attack based on local low-frequency perturbation. Specifically, instead of imposing perturbation on the entire audio sample, a local attacking region is firstly sought, confining the perturbed distortion to a local region. Second, considering that the majority of energy concentrates on the low-frequency bands, the proposed method suggests performing perturbation generation in the low-frequency domain. Experimental results demonstrate that, compared with the recent methods, our method could implement target attacking to SRS with a higher attacking success rate, at the cost of much lower queries and adversarial perturbation. Jiacheng Deng 0001, Li Dong 0006, Rangding Wang, Rui Yang 0006, Diqun Yan |
IEEE Signal Process. Lett. | 1 |