EDBT 2026 Demo / reviewers in the wild / expert
Diqun Yan
dblp:52/8076
· DBLP profile ↗
58ranked-venue papers
9as first author
43since 2021 · last 2026
0000-0002-5241-7276ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Security and privacy · 13 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive multi-task adversarial attacks on audio tasks
Jiazhen Jia, Rangding Wang, Diqun Yan |
Expert Syst. Appl. | 4 |
| 2026 | Amplifying discriminative distortions: A generative latent feature reinforcement framework for audio spoofing detection
Site Wu, Zhe Ye 0001, Rangding Wang, Diqun Yan |
Expert Syst. Appl. | 6 |
| 2026 | Prompt-based contrastive learning for non-intrusive speech quality assessment
Diqun Yan, Naiyuan Li |
Speech Commun. | 1 |
| 2026 | DUAP: Disentanglement-Based Universal Adversarial Perturbations for Robust Multilingual Speech Privacy ProtectionabstractThe rapid advancement of automatic speech recognition (ASR) models has significantly bolstered their multilingual proficiency and robustness, amplifying concerns over user speech privacy. Attackers may use hidden microphones or network attacks to capture and transcribe sensitive user interactions. Whisper, a state-of-the-art (SOTA) multilingual speech recognition model, delivers exceptional transcription accuracy across diverse languages. However, its superior performance also extends privacy leakage risks to multilingual contexts. Previous privacy-preserving methods based on adversarial examples were primarily optimized for monolingual models, limiting their effectiveness in multilingual settings. Moreover, as these perturbation mechanisms were predominantly tailored for English, their transferability to other languages remains constrained. To address this vulnerability, we propose the Disentanglement-based Universal Adversarial Perturbation (DUAP), a privacy-preserving method designed to counteract the Whisper model. Unlike optimization-based approaches, DUAP embeds language-specific features in the latent space to generate robust adversarial perturbations, providing consistent protection across multiple languages and effectively mitigating privacy risks in multilingual contexts. The method employs a two-stage language attack: first, a Language Feature Disentanglement model disentangles and reconstructs language-specific features to produce adversarial examples (AEs); second, gradient-based optimization refines AEs to disrupt Whisper’s language identification module. DUAP’s perturbations, effective in physical and digital settings, achieve SNRs from 40 dB (lightest) to above 17 dB (strongest). Across three Whisper model sizes, DUAP yields WERs over 95% (English), 85% (other languages), and 87% (physical settings), maintaining above 96% under AAC (64, 72 kbps) and MP3 (32, 96 kbps) compressions. Jiazhen Jia, Rangding Wang, Diqun Yan |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Speed Master: Quick or Slow Play to Attack Speaker RecognitionabstractBackdoor attacks pose a significant threat during the model's training phase. Attackers craft pre-defined triggers to break deep neural networks, ensuring the model accurately classifies clean samples during inference yet erroneously classifies samples added with these triggers. Recent studies have shown that speaker recognition systems trained on large-scale data are susceptible to backdoor attacks. Existing attackers employ unnoticed ambient sounds as triggers. However, these sounds are not inherently part of the training samples themselves. In essence, triggers can be designed to maintain an intrinsic connection with the original speech to enhance stealthiness. Our paper presents a novel attack methodology named Speed Master, which undermines deep neural networks by manipulating the speed of speech samples. Specifically, we execute poison-only backdoor attacks using speed or tempo adjustment. Changes in speech rate have become a common occurrence, as seen on platforms that allow users to adjust playback speed. In real-world scenarios, people naturally adjust their speaking rate depending on the context. As a result, changes in a speaker’s speech rate are typically perceived as normal and are unlikely to raise suspicion. Furthermore, detecting such subtle adjustments becomes challenging for users without reference speech. Our comprehensive experiments demonstrate that Speed Master can achieve an ASR over 99% in the digital domain, with only a 0.6% poisoning rate. Additionally, we validate the feasibility of Speed Master in the real world and its resistance to typical defensive measures. Zhe Ye 0001, Ying Ren, Xiangui Kang, Diqun Yan, Bin Ma 0003, Shiqi Wang 0001 |
AAAI | 5 |
| 2025 | DyMEvalNet: Dynamic Text-Audio-Personalization Fusion for Multimodal Music Quality AssessmentabstractText-to-music (TTM) generation has gained significant attention for its ability to produce customized music from textual prompts. However, current methods lack objective standards and semantic alignment metrics for evaluating generated music quality. To address this issue, we propose DyMEval-Net, a dynamic multimodal music quality assessment model based on a novel three-branch architecture that incorporates listener personalization into TTM evaluation for the first time. DyMEvalNet integrates audio features, text embeddings, and personalized listener embeddings to jointly process audio-text semantics. Evaluated on the MusicEval dataset, DyMEvalNet achieves state-of-the-art performance, improving utterance-level Spearman Rank Correlation Coefficient (SRCC) scores by 19.97% for overall music quality (ranking second in the AudioMOS Challenge 2025) and 18.27% for textual alignment. The proposed model demonstrates robust effectiveness in assessing TTM-generated music quality by simultaneously considering audio fidelity and semantic alignment. Xiaoxun Wu, Kailai Shen, Naiyuan Li, Diqun Yan |
ASRU | 5 |
| 2025 | SML: A Backdoor Defense for Non-Intrusive Speech Quality Assessment via Semi-Supervised and Multi-Task LearningabstractNon-intrusive speech quality assessment (NISQA) is widely used in speech downstream tasks due to its ability to predict the quality of speech without a reference speech. However, few researchers have focused on the backdoor security of NISQA. Despite the backdoor defenses have been extensively studied to mitigate the threat of maliciously modifications in deep neural networks. In particular, semi-supervised based backdoor defenses have excellent defensive performance by depriving backdoor attacks of their most essential need. But these defense methods rely on data-augmentation consistency and thus cannot be applied to NISQA. In this work, we propose a backdoor defense based on semi-supervised and multi-task learning (SML). Semi-supervised learning is based on the simple assumption that the same input should be as consistent as possible in two similar models. Multi-task learning further improves the prediction performance of mean opinion score (MOS) by learning the tasks of perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI) and speech distortion index (SDI). Extensive experiments involving five backdoor defenses against five backdoor attacks on two benchmark datasets demonstrate the superiority of our SML approach. Ying Ren, Jiahong Ye, Diqun Yan, Bin Ma 0003 |
ICASSP | 5 |
| 2025 | CBA: Backdoor Attack on Deep Speech Classification via Audio Compression
Ying Ren, Diqun Yan |
INTERSPEECH | 4 |
| 2025 | Diversity-Preserving Robust Watermarking for Diffusion Model Generated ImagesabstractThis paper introduces a robust watermarking technique for diffusion model-generated images, which effectively balances watermark robustness, image fidelity, and diversity preservation. Unlike traditional post-hoc approaches, the proposed method embeds watermark information directly into the latent noise of the diffusion model, ensuring seamless integration into the image generation process. This approach minimizes perceptual impact while maintaining high visual quality and diversity of the generated images. Experimental results demonstrate the method’s resilience to various image distortions, including noise, compression, blurring etc., significantly outperforming existing watermarking techniques. The proposed method supports both watermark detection and bit-level extraction, providing a practical solution for secure content protection and traceability in generative models without compromising the integrity of the image generation process. Linghong Wan, Li Dong 0006, Diqun Yan, Rangding Wang |
ISCAS | 3 |
| 2025 | Exploiting Hard Samples for Stealthy Backdoor Attacks on Large Language Models
Diqun Yan, Rangding Wang |
NPC (2) | 1 |
| 2025 | Improving crowdsourced label quality by peer-to-peer federated learning
Xiangming Lu, Jiangbo Qian, Chong Wang 0001, Diqun Yan, Youhui Zhang |
Appl. Intell. | 4 |
| 2025 | An efficient and scalable semi-supervised framework for semantic segmentation
Huazheng Hao, Hui Xiao 0005, Li Dong 0006, Diqun Yan, Dongtai Liang, Jiayan Zhuang, Chengbin Peng 0001 |
Neural Comput. Appl. | 5 |
| 2025 | One-class network leveraging spectro-temporal features for generalized synthetic speech detection
Jiahong Ye, Diqun Yan, Songyin Fu, Bin Ma 0003, Zhihua Xia |
Speech Commun. | 2 |
| 2025 | Paradoxical Role of Adversarial Attacks: Enabling Crosslinguistic Attacks and Information Hiding in Multilingual Speech RecognitionabstractWith the rise of automatic speech recognition (ASR) research and practical applications, enabling adversarial attacks on ASR systems via subtle perturbations has become a priority. Most prior research has focused on single-language, single-model ASR systems. However, multilingual ASR systems hold opportunities for crosslinguistic attacks and covert message transmission. This letter introduces a new approach for crosslinguistic adversarial attacks in multilingual ASR, focusing on information hiding. For example, in military settings, adversarial examples applied to eavesdropping devices can encode messages detectable only by friendly devices, leaving adversaries, even with identical methods, unable to access them. This letter examines multilingual ASR system properties and introduces a crosslinguistic adversarial example with minimal perturbation, allowing friendly classifiers to extract hidden information while being undetectable by hostile classifiers. The experimental results on 5 models and 5 datasets show that the proposed method achieves a success rate of over 90% and an SNR close to 40 dB. Zhihua Xia, Bin Ma 0003, Diqun Yan |
IEEE Signal Process. Lett. | 4 |
| 2025 | Facial Data Minimization: Shallow Model as Your Privacy FilterabstractFace recognition service has been widely adopted across various domains, offering significant convenience and enhancing efficiency in numerous applications. However, once a user's facial data is transmitted to a service provider, the user will lose control over his/her biometric data. In recent years, there have been various security and privacy issues due to the leakage of facial data. Although many privacy enhancement methods have been proposed, they usually fail when they are not accessible to adversaries' strategies or the complete face recognition model. Therefore, in this work, we propose a Privacy Minimization Transformation (PMT) method, designed to address two common scenarios in practical face recognition systems: the uploading of facial images and facial features. This method can process the private facial data based on the shallow network of the face recognition model to obtain the obfuscated data. The obfuscated data cannot only maintain satisfactory performance on the authorized models (i.e., the models specified by the user) and restrict the performance on other unauthorized models (i.e., the models not specified by the user) but also prevent privacy data from leaking by AI methods and human visual theft. Additionally, since a service provider may execute preprocessing operations on the received data, we propose an enhanced perturbation method to improve the robustness of PMT. Besides, to authorize one facial image to multiple service models simultaneously, a multiple-restriction mechanism is proposed to improve the scalability of PMT. Finally, we conduct extensive experiments and evaluate the effectiveness of the proposed PMT against face reconstruction, function creep, and face attribute estimation attacks. Experimental results demonstrate that PMT performs well in preventing facial function creep and privacy leakage while maintaining high face recognition accuracy Yuwen Pu, Jiayu Pan, Diqun Yan, Xuhong Zhang 0002, Shouling Ji |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Breaking Speaker Recognition with PaddingbackabstractMachine Learning as a Service (MLaaS) has gained popularity due to advancements in Deep Neural Networks (DNNs). However, untrusted third-party platforms have raised concerns about AI security, particularly in backdoor attacks. Recent research has shown that speech backdoors can utilize transformations as triggers, similar to image backdoors. However, human ears can easily be aware of these transformations, leading to suspicion. In this paper, we propose PaddingBack, an inaudible backdoor attack that utilizes malicious operations to generate poisoned samples, rendering them indistinguishable from clean ones. Instead of using external perturbations as triggers, we exploit the widely-used speech signal operation, padding, to break speaker recognition systems. Experimental results demonstrate the effectiveness of our method, achieving a significant attack success rate while retaining benign accuracy. Furthermore, Padding-Back demonstrates the ability to resist defense methods and maintain its stealthiness against human perception. Zhe Ye 0001, Diqun Yan, Li Dong 0006, Kailai Shen |
ICASSP | 2 |
| 2024 | EventTrojan: Manipulating Non-Intrusive Speech Quality Assessment via Imperceptible EventsabstractNon-Intrusive speech quality assessment (NISQA) has gained significant attention for predicting speech’s mean opinion score (MOS) without requiring the reference speech. Researchers have gradually started to apply NISQA to various practical scenarios. However, little attention has been paid to the security of NISQA models. Backdoor attacks represent the most serious threat to deep neural networks (DNNs) due to the fact that backdoors possess a very high attack success rate once embedded. However, existing backdoor attacks assume that the attacker actively feeds samples containing triggers into the model during the inference phase. This is not adapted to the specific scenario of NISQA. And current backdoor attacks on regression tasks lack an objective metric to measure the attack performance. To address these issues, we propose a novel backdoor triggering approach (EventTrojan) that utilizes an event during the usage of the NISQA model as a trigger. Moreover, we innovatively provide an objective metric for backdoor attacks on regression tasks. Extensive experiments on four benchmark datasets demonstrate the effectiveness of the EventTrojan attack. Besides, it also has good resistance to several defense methods. Ying Ren, Kailai Shen, Zhe Ye 0001, Diqun Yan |
ICME | 4 |
| 2024 | Non-intrusive speech quality assessment: A survey
Kailai Shen, Diqun Yan, Zhe Ye 0001 |
Neurocomputing | 2 |
| 2024 | C²F²: Cross-Task Cross-Domain Feature Fusion for Semi-Supervised Change DetectionabstractSemi-supervised learning for change detection (CD), which significantly reduces the labor costs associated with data annotation, has recently garnered substantial attention. In this study, we propose to enhance traditional semi-supervised learning frameworks by leveraging cross-task cross-domain (CTCD) models, which generate complementary features that differ from standard hidden features. The procedure is as follows. First, the standard features obtained from a traditional encoding–decoding structure are fused with attention-augmented complementary features. Second, a secondary decoder maps the fused heterogeneous features into the label space to obtain high-quality pseudo-labels, offering more precise guidance for semi-supervised learning on traditional structures. This approach improves pseudo-labels by leveraging the strength of CTCD models, including large pretrained models, to enhance the semi-supervised learning process of domain-specific and task-specific models. Experimental results on benchmark datasets demonstrate that our proposed approach surpasses state-of-the-art methods. Dongjie Zhang 0001, Yuting Hong, Xiaojie Qiu, Li Dong 0006, Diqun Yan, Chengbin Peng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Mixed-Bit Sampling Graphic: When Watermarking Meets Copy Detection PatternabstractCopy Detection Pattern (CDP) is a high-density random noise-alike image that exhibits a different noise pattern after physical copying, and is thus treated as a promising anti-counterfeiting solution. However, CDP cannot convey any message, and it is often used in combination with additional carriers, such as QR codes. In this letter, we take the first step towards extending CDP with watermarking functionality. Specifically, we devise a scheme called Mixed-bit Sampling Graphic (MSG), which could realize invisible watermarking and anti-counterfeiting simultaneously. Compared with conventional CDP, the noise pattern generation of MSG is controlled by the portions of sampling over two bit templates. We formulate this mixed-bit sampling process as an optimization problem and solve it using a block coordinate descent sampling algorithm. Experimental results validate that the proposed MSG can effectively communicate watermark bits while retaining the anti-counterfeiting capability of CDP. Li Dong 0006, Rangding Wang, Diqun Yan, Chengbin Peng 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Dynamic Sensing and Correlation Loss Detector for Small Object Detection in Remote Sensing ImagesabstractRecently, significant object detection achievements have been emerged for optical remote sensing images. However, the performance and efficiency of small object detection are still highly unsatisfactory because of the scale diversity between the objects; furthermore, small objects always have small amounts of effective information that are difficult to locate. To address this problem, we propose a novel dynamic sensing and correlation loss detector (DCDet) for performing object detection in remote sensing images. The detector consists of two modules: a small-object dynamic sensing (SODS) module and a simple but effective correlation loss function (CrLoss). SODS is utilized to capture the information of small objects in a scale sequence. We consider the feature pyramid as a set of video frames when the camera is zoomed in on the image and use the object focusing module in dynamic sensing to always focus on the small objects in each video frame. The detection performance achieved for small objects is improved by shifting the detector’s attention from the entire image to small objects within the frame to provide a multiscale feature representation of the small objects and their contextual information. The CrLoss is a special correlation loss for remote sensing image object detection tasks and directly optimizes the correlation coefficient to improve the performance of a detector. Extensive experiments conducted on the publicly available DOTA, DIOR-R and HRSC2016 datasets show that our DCDet outperforms the existing state-of-the-art remote sensing object detection methods in terms of many evaluation metrics. Chongchong Shen, Jiangbo Qian, Chong Wang 0001, Diqun Yan, Caiming Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Uncertainty-Guided Contrastive Learning for Weakly Supervised Point Cloud SegmentationabstractThree-dimensional point cloud data are widely used in many fields, as they can be easily obtained and contain rich semantic information. Recently, weakly supervised segmentation has attracted lots of attention, because it only requires very few labels, thus reducing time-consuming and expensive data annotation efforts for huge amounts of point cloud data. The existing approaches typically adopt softmax scores from the last layer as the confidence for selecting high-confident point predictions. However, such approaches can ignore the potential value of a large number of low-confidence point predictions under traditional metrics. In this work, we propose an uncertainty-guided contrastive learning (UCL) framework for weakly supervised point cloud segmentation. A novel uncertainty metric based on prototype entropy (PE) is presented to estimate the reliability of model predictions. With this metric, we propose a negative contrastive learning module exploiting negative pseudo-labels of predictions with low reliability and an active contrastive learning module enhancing feature learning of segmentation models by predictions with high reliability. We also propose a generic multiscale feature perturbation method to expand a wider perturbation space. Extensive experimental results on indoor and outdoor point cloud datasets demonstrate that the proposed method achieves competitive performance. Baochen Yao, Li Dong 0006, Xiaojie Qiu, Kangkang Song, Diqun Yan, Chengbin Peng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Multi-Level Label Correction by Distilling Proximate Patterns for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation relieves the reliance on large-scale labeled data by leveraging unlabeled data. Recent semi-supervised semantic segmentation approaches mainly resort to pseudo-labeling methods to exploit unlabeled data. However, unreliable pseudo-labeling can undermine the semi-supervision processes. In this paper, we propose an algorithm called Multi-Level Label Correction (MLLC), which aims to use graph neural networks to capture structural relationships in Semantic-Level Graphs (SLGs) and Class-Level Graphs (CLGs) to rectify erroneous pseudo-labels. Specifically, SLGs represent semantic affinities between pairs of pixel features, and CLGs describe classification consistencies between pairs of pixel labels. With the support of proximate pattern information from graphs, MLLC can rectify incorrectly predicted pseudo-labels and can facilitate discriminative feature representations. We design an end-to-end network to train and perform this effective label corrections mechanism. Experiments demonstrate that MLLC can significantly improve supervised baselines and outperforms state-of-the-art approaches in different scenarios on Cityscapes and PASCAL VOC 2012 datasets. Specifically, MLLC improves the supervised baseline by at least 5% and 2% with DeepLabV2 and DeepLabV3+ respectively under different partition protocols. Hui Xiao 0005, Yuting Hong, Li Dong 0006, Diqun Yan, Jiayan Zhuang, Dongtai Liang, Chengbin Peng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | SQAT-LD: SPeech Quality Assessment Transformer Utilizing Listener Dependent Modeling for Zero-Shot Out-of-Domain MOS PredictionabstractIn this paper, we propose the speech quality assessment transformer utilizing listener dependent modeling (SQAT-LD) mean opinion score (MOS) prediction system, which was submitted to the 2023 VoiceMOS Challenge. The system is based on a combination of self-supervised learning (SSL) models and listener-dependent modeling. Due to this challenge’s emphasis on real-world and challenging zero-shot out-of-domain MOS prediction in three different voice evaluation scenarios, we specifically designed a two-branch module to predict scores and weights for each frame, aiming to achieve better generalization. In the challenge, our system achieved fourth place in Track 1a, second place in Track 1b and first place in Track 2. Additionally, we conducted an ablation study to investigate the effectiveness of our proposed method. Kailai Shen, Diqun Yan, Li Dong 0006, Ying Ren, Xiaoxun Wu |
ASRU | 2 |
| 2023 | Animal Re-Identification Algorithm for Posture DiversityabstractRe-identification (Re-ID) technology is important for wildlife conservation and intelligent farm management. With the development of deep learning, the performance of animal ReID based on computer vision has been improved. However, variations in animal pose push a negative impact on recognition performance. In this paper, a Multi-pose Feature Fusion Network (MPFNet) is proposed to improve the performance of the Re-ID. First, we construct three pose modules for the three postures, that is, standing, sitting, and lying, respectively. In each pose module, there are two parallel branches, one is a global branch for extracting global features, and the other is a local branch for extracting local features. In addition, to obtain more effective feature representations, we weighted fusion for the global branching of the three pose modules. We validate the efficiency of MPFNet on both the self-built MPDD dog dataset and the public ATRW Amur Tiger dataset. Experimental results show that MPFNet can obtain better recognition performance than other state-of-the-art Re-ID methods. The source of code will be public available at https://github.com/hezhimin7028/MPFNet. Jiangbo Qian, Diqun Yan, Chong Wang 0001 |
ICASSP | 3 |
| 2023 | A Pseudo-Dual Self-Rectification Framework for Semantic SegmentationabstractSemantic segmentation has achieved remarkable success in various applications. However, the training process for such techniques necessitates a significant amount of labeled data. Although semi-supervised frameworks can alleviate this issue, traditional approaches typically require multiple baseline models to form a dual model. To allow a semi-supervised semantic segmentation framework to be used in robotic systems with precious computation and memory resources, we propose a framework utilizing a single baseline model only. The overall framework is composed of three parts: an encoder, a shallow decoder, and a deep decoder. It distills knowledge from the ensemble of two decoders to improve the encoder, which can implicitly form a pseudo-dual model. It also calculates class-wise likelihoods according to the similarity between features and class prototypes learned from different decoders and rectifies low-confidence pseudo-labels. Our framework outperforms state-of-the-art frameworks on benchmark datasets with a significant amount of decrease in using computing resources. Huazheng Hao, Hui Xiao 0005, Li Dong 0006, Diqun Yan, Dongtai Liang, Jiayan Zhuang, Chengbin Peng 0001 |
ICME | 4 |
| 2023 | Gradient Sign Inversion: Making an Adversarial Attack a Good DefenseabstractDeep neural networks have been proven vulnerable to deliberately crafted adversarial example, which cause serious safety and security concerns. Many defense approaches were proposed to resist such threats. However, existing defenses such as pre-compression or adversarial training would degrade the model performance on clean images or incur heavy computational costs. In this work, we propose a plug-and-play defensive module Gradient Sign Inversion (GSI) to defend gradient-based attack. Essentially, GSI attempts to inverse the direction of the backpropagated gradient for the victim model, disturbing the adversarial example generation of the attacking while retaining the performance of the vanilla network on genuine inputs. Specifically, an additive model based on periodic trigonometric function is established by investigating the necessary conditions that a suitable defensive module should have. By enforcing constraints on the defensive module, the parameters of GSI are determined, accompanied by a theoretical justification. Interestingly, we observe that the proposed GSI not only prevents the gradient-based adversarial attack, but can even improve the confidence of the ground-truth label when initiating an attack, making the attack betray as a defense. Source code is publicly available at https://github.com/JidaDiao/GSI. Xiaojian Ji, Li Dong 0006, Rangding Wang, Diqun Yan, Yang Yin, Jinyu Tian 0001 |
IJCNN | 4 |
| 2023 | Fake the Real: Backdoor Attack on Deep Speech Classification via Voice ConversionabstractDeep speech classification has achieved tremendous success and greatly promoted the emergence of many real-world applications. However, backdoor attacks present a new security threat to it, particularly with untrustworthy third-party platforms, as pre-defined triggers set by the attacker can activate the backdoor. Most of the triggers in existing speech backdoor attacks are sample-agnostic, and even if the triggers are designed to be unnoticeable, they can still be audible. This work explores a backdoor attack that utilizes sample-specific triggers based on voice conversion. Specifically, we adopt a pre-trained voice conversion model to generate the trigger, ensuring that the poisoned samples does not introduce any additional audible noise. Extensive experiments on two speech classification tasks demonstrate the effectiveness of our attack. Furthermore, we analyzed the specific scenarios that activated the proposed backdoor and verified its resistance against fine-tuning. Zhe Ye 0001, Terui Mao, Li Dong 0006, Diqun Yan |
INTERSPEECH | 4 |
| 2023 | Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character RecognitionabstractOptical Character Recognition (OCR) enables automatic text extraction from scanned or digitized text images, but it also makes it easy to pirate valuable or sensitive text from these images. Previous methods to prevent OCR piracy by distorting characters in text images are impractical in real-world scenarios, as pirates can capture arbitrary portions of the text images, rendering the defenses ineffective. In this work, we propose a novel and effective defense mechanism termed the Universal Defensive Underpainting Patch (UDUP) that modifies the underpainting of text images instead of the characters. UDUP is created through an iterative optimization process to craft a small, fixed-size defensive patch that can generate non-overlapping underpainting for text images of any size. Experimental results show that UDUP effectively defends against unauthorized OCR under the setting of any screenshot range or complex image background. It is agnostic to the content, size, colors, and languages of characters, and is robust to typical image operations such as scaling and compressing. In addition, the transferability of UDUP is demonstrated by evading several off-the-shelf OCRs. The code is available at https://github.com/QRICKDD/UDUP. Jiacheng Deng 0001, Li Dong 0006, Diqun Yan, Rangding Wang, Dengpan Ye, Lingchen Zhao, Jinyu Tian 0001 |
ACM Multimedia | 4 |
| 2023 | Adaptive-SpEx: Local and Global Perceptual Modeling with Speaker Adaptation for Target Speaker ExtractionabstractTarget speaker extraction aims to extract a target speaker's speech from a multi-talker environment with the help of the target speaker's reference speech. However, the simple fusion of different features and local perceptual modeling lead to limited extraction performance. In this work, we propose a new speaker extraction model called Adaptive-SpEx. The correlation between mixed speech features and speaker embedding is fully exploited, and a dual-path structure is used for local and global perceptual modeling. We evaluate the model on the WSJ0-2mix-extr dataset in terms of its ability to reconstruct signal quality. Experimental results show that the proposed model outperforms other baseline systems on WSJ0-2mix-extr and achieves better generalizability on the Libri-2talker dataset. Furthermore, the proposed model can significantly reduce the word error rate of mixed speech on speech recognition from 79.49% to 32.73%. Xianbo Xu, Diqun Yan, Li Dong 0006 |
SMC | 2 |
| 2023 | Corrigendum to "Semi-supervised semantic segmentation with cross teacher training" [Neurocomputing 508 (2022) 36-46]
Hui Xiao 0005, Li Dong 0006, Shuibo Fu, Diqun Yan, Kangkang Song, Chengbin Peng 0001 |
Neurocomputing | 5 |
| 2023 | Imperceptible adversarial audio steganography based on psychoacoustic model
Lang Chen, Rangding Wang, Li Dong 0006, Diqun Yan |
Multim. Tools Appl. | 4 |
| 2023 | Stealthy Backdoor Attack Against Speaker Recognition Using Phase-Injection Hidden TriggerabstractDeep learning has achieved significant breakthroughs in speaker recognition, driven by continual advancements in foundation models. However, malicious third-party platforms have introduced a severe security concern through backdoor attacks, in which attackers can manipulate a model to output a specific label by implanting a trigger. Existing speech backdoor attack methods typically utilize fixed and unnoticeable perturbations as triggers, but these may still be audible and thus detected during training and inference stages. To overcome this limitation, we propose a novel backdoor attack paradigm (PhaseBack) injecting triggers in the phase spectrum. PhaseBack exhibits sufficient stealth by leveraging the fact that the human ear is insensitive to phase information. Besides, injecting partial perturbations in the frequency domain results in global perturbations throughout the time domain, making the attack more effective. Extensive experiments on the Voxceleb1 dataset demonstrate the effectiveness and stealthiness of PhaseBack. Moreover, it has strong resistance to bypass several defense methods. Zhe Ye 0001, Diqun Yan, Li Dong 0006, Jiacheng Deng 0001, Shui Yu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Physical Anti-copying Semi-robust Random Watermarking for QR Code
Li Dong 0006, Rangding Wang, Diqun Yan, Weiwei Sun 0009, Hang-Yu Fan |
IWDW | 4 |
| 2022 | On Attacking Deep Image Quality Evaluator Via Spatial TransformabstractAdversarial examples fool the neural networks by adding slightly-perturbed noise to the original image, which barriers the usability of deep models. Most of the works focused on the adversarial attack on the classification task. We, in this work, attempt to develop an adversarial example generation method for attacking neural-based image quality assessment (IQA). Specifically, instead of employing conventional additive adversarial noise generation methods, we propose an image content deformation approach, avoiding the loss of adversarial noise after compression. The deformation component is designed as neural layers. The given image is firstly deformed and then undergoes compression; an existing IQA evaluates the compressed image. The deformation layers are trained by back-propagating the differences between the targeted IQA score and the originally-evaluated one. Experimental results demonstrate that the proposed method can produce compression-resistant adversarial images for image quality evaluators. The generated adversarial examples could effectively attack neural-based image quality evaluators with less distortion. The data and code of this work are available at https://github.com/luning409/Attack_IQA. Li Dong 0006, Diqun Yan, Xianliang Jiang |
SMC | 3 |
| 2022 | Robust Document Image Forgery Localization Against Image BlendingabstractDigital documents, as a twin of hard copy, are increasingly being used as credible evidence. Unfortunately, digital document images easily suffer forgery or malicious manipulation, with the availability of sophisticated image editing tools. To verify and detect the possible forgeries for a given document, a number of forensic schemes have been developed. However, in the real-world scenario, the doctored image could be further processed or transmitted over a channel with unknown distortion, which dramatically degrade the forgery detection performance. In this work, we make the first step towards designing a robust document image forgery localization against image blending. Specifically, we propose an encoder-decoder neural network architecture consisting of three modules. The first module is responsible for capturing the multi-scale features from the high-level feature maps, and the remaining two attention-based modules aim to extract low-level local features and high-level global features. For training the model, we construct a dedicated forgery document database processed by several recent image blending procedures. Extensive experiments demonstrate the effectiveness and superiority of the proposed method in detecting the forgery that undergoes image blending. The source code, models and the constructed image dataset are publicly available at https://github.com/lwp0201/Image-Forgery-Localization-Against-Image-Blending. Weipeng Liang, Li Dong 0006, Rangding Wang, Diqun Yan, Yuanman Li |
TrustCom | 4 |
| 2022 | Semi-supervised semantic segmentation with cross teacher training
Hui Xiao 0005, Li Dong 0006, Shuibo Fu, Diqun Yan, Kangkang Song, Chengbin Peng 0001 |
Neurocomputing | 5 |
| 2022 | Anti-forensics of fake stereo audio using generative adversarial network
Tianyun Liu, Diqun Yan |
Multim. Tools Appl. | 2 |
| 2022 | Decision-Based Attack to Speaker Recognition System via Local Low-Frequency PerturbationabstractDespite neural network-based speaker recognition systems (SRS) have enjoyed significant success, they are proved to be quite vulnerable to adversarial examples. In practice, the SRS model parameters are not always available. Attackers have to probe the model only via querying, and such decision-based attacking merely relies on the output label is quite challenging. This letter proposes a two-step query-efficient decision-based attack based on local low-frequency perturbation. Specifically, instead of imposing perturbation on the entire audio sample, a local attacking region is firstly sought, confining the perturbed distortion to a local region. Second, considering that the majority of energy concentrates on the low-frequency bands, the proposed method suggests performing perturbation generation in the low-frequency domain. Experimental results demonstrate that, compared with the recent methods, our method could implement target attacking to SRS with a higher attacking success rate, at the cost of much lower queries and adversarial perturbation. Jiacheng Deng 0001, Li Dong 0006, Rangding Wang, Rui Yang 0006, Diqun Yan |
IEEE Signal Process. Lett. | 5 |
| 2021 | Fast speech adversarial example generation for keyword spotting system with conditional GAN
Donghua Wang 0001, Li Dong 0006, Rangding Wang, Diqun Yan |
Comput. Commun. | 4 |
| 2021 | Exposing Speech Transsplicing Forgery with Noise Level InconsistencyabstractSplicing is one of the most common tampering techniques for speech forgery in many forensic scenarios. Some successful approaches have been presented for detecting speech splicing when the splicing segments have different signal-to-noise ratios (SNRs). However, when the SNRs between the spliced segments are close or even same, no effective detection methods have been reported yet. In this study, noise inconsistency between the original speech and the inserted segment from other speech is utilized to detect the splicing trace. First, noise signal of the suspected speech is extracted by a parameter-optimized noise estimation algorithm. Second, the statistical Mel frequency features are extracted from the estimated noise signal. Finally, the spliced region is located by utilizing a change point detection algorithm on the estimated noise signal. The effectiveness of the proposed method is evaluated on a well-designed speech splicing dataset. The comparative experimental results show that the proposed algorithm can achieve better detection performance than other algorithms. Diqun Yan, Mingyu Dong, Jinxing Gao |
Secur. Commun. Networks | 1 |
| 2021 | Tackling the Cover Source Mismatch Problem in Audio Steganalysis With Unsupervised Domain AdaptationabstractNowadays, the convolutional neural network (CNN) based steganalysis has achieved remarkable performance in the well-controlled lab environment. However, the cover source mismatch (CSM) problem, which can be attributed to the discrepancy between the training, and evaluation datasets, is still one of the pivotal obstacles for adapting the steganalysis into real-world applications. In this letter, we propose to merge the domain adaptation strategy into CNN-based audio steganalysis for handling the CSM problem. Specifically, the proposed framework contains three components: feature extractor, steganalytic classifier, and domain discriminator. The cascade of feature extractor, and steganalytic classifier compose the typical supervised steganalysis model. The unsupervised domain adaptation is implemented by the domain adversarial training between the feature extractor, and domain discriminator. Ultimately, the feature extractor is trained to extract the steganalytic, and domain-invariant features. It aims to reduce the domain gap between the training data, and testing data. The experimental results show that our approach could effectively mitigate the CSM impact caused by the diversity of audio recording devices. Yuzhen Lin, Rangding Wang, Li Dong 0006, Diqun Yan, Jie Wang 0028 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Antiforensics of Speech Resampling Using Dual-Path StrategyabstractResampling is an operation to convert a digital speech from a given sampling rate to a different one. It can be used to interface two systems with different sampling rates. Unfortunately, resampling may also be intentionally utilized as a postoperation to remove the manipulated artifacts left by pitch shifting, splicing, etc. To detect the resampling, some forensic detectors have been proposed. Little consideration, however, has been given to the security of these detectors themselves. To expose weaknesses of these resampling detectors and hide the resampling artifacts, a dual‐path resampling antiforensic framework is proposed in this paper. In the proposed framework, 1D median filtering is utilized to destroy the linear correlation between the adjacent speech samples introduced by resampling on low‐frequency component. And for high‐frequency component, Gaussian white noise perturbation (GWNP) is adopted to destroy the periodic resampling traces. The experimental results show that the proposed method successfully deceives the existing resampling forensic algorithms while keeping good perceptual quality of the resampled speech. Diqun Yan, Yongkang Gong 0002, Tianyun Liu |
Wirel. Commun. Mob. Comput. | 1 |
| 2020 | Towards Designing an Effective Complexity Indicator for Audio SteganographyabstractIn the field of steganography, to effectively hide the secret message, it is of great importance to determine which part of the steganographic cover is suitable for embedding. Currently, most of the existing works focus on the image cover, while few works touch the audio cover case. In this work, we attempt to characterize the complexity of audio for selecting the steganographic cover. Specifically, the original cover is first convoluted with a specially designed adaptive convolution kernel. Based on the residual between the original and the convoluted audio, we derive a quantity for measuring the complexity of each frame for a given audio clip. Experimental results verify the usability of the proposed complexity indicator, suggesting high-complexity audio cover is favorable for data embedding. It is also found that the proposed complexity indicator could further boost the steganographic performance of the state-of the-art audio steganography methods. The source code is publicly available at https://github.com/capzxy/audio-complexity. Xueyuan Zhang, Rangding Wang, Li Dong 0006, Diqun Yan, Yuzhen Lin, Jie Wang 0028 |
ICC | 4 |
| 2020 | Efficient Generation of Speech Adversarial Examples with Generative Model
Donghua Wang 0001, Rangding Wang, Li Dong 0006, Diqun Yan |
IWDW | 4 |
| 2020 | An Antiforensic Method against AMR Compression DetectionabstractAdaptive multirate (AMR) compression audio has been exploited as an effective forensic evidence to justify audio authenticity. Little consideration has been given, however, to antiforensic techniques capable of fooling AMR compression forensic algorithms. In this paper, we present an antiforensic method based on generative adversarial network (GAN) to attack AMR compression detectors. The GAN framework is utilized to modify double AMR compressed audio to have the underlying statistics of single compressed one. Three state-of-the-art detectors of AMR compression are selected as the targets to be attacked. The experimental results demonstrate that the proposed method is capable of removing the forensically detectable artifacts of AMR compression under various ratios with an average successful attack rate about 94.75%, which means the modified audios generated by our well-trained generator can treat the forensic detector effectively. Moreover, we show that the perceptual quality of the generated AMR audio is well preserved. Diqun Yan, Li Dong 0006, Rangding Wang |
Secur. Commun. Networks | 1 |
| 2019 | Audio Steganalysis with Improved Convolutional Neural NetworkabstractDeep learning, especially the convolutional neural network (CNN), has enjoyed significant success in many fields, e.g., image recognition. Recently, CNN has successfully applied to multimedia steganalysis. However, the detection performance is still unsatisfactory. In this work, we propose an improved CNN-based method for audio steganalysis. Specifically, a special convolutional layer is first carefully designed, which could capture the minor steganographic noise. Then, a truncated linear unit is adapted to activate the output of shallow convolutional layer. In addition, we employ the average pooling to minimize the over-fitting risk. Finally, a parameter transfer strategy is adopted, aiming to boost the detection performance for the low embedding-rate cases. The experimental results evaluated on 30,000 audio clips verify the effectiveness of our method for a variety of embedding rates. Compared with the existing CNN-based steganalysis methods, our proposed method could achieve superior performance. To facilitate the reproducible research, the source code will be released at GitHub. Yuzhen Lin, Rangding Wang, Diqun Yan, Li Dong 0006, Xueyuan Zhang |
IH&MMSec | 3 |
| 2017 | Steganalysis of MP3Stego with low embedding-rate using Markov feature
Chao Jin 0003, Rangding Wang, Diqun Yan |
Multim. Tools Appl. | 3 |
| 2016 | Source Cell-Phone Identification Using Spectral Features of Device Self-noise
Chao Jin 0003, Rangding Wang, Diqun Yan, Biaoli Tao, Anshan Pei |
IWDW | 3 |
| 2016 | An efficient algorithm for double compressed AAC audio detection
Chao Jin 0003, Rangding Wang, Diqun Yan, Jinglei Zhou |
Multim. Tools Appl. | 3 |
| 2015 | Multiple MP3 Compression Detection Based on the Statistical Properties of Scale Factors
Jinglei Zhou, Rangding Wang, Chao Jin 0003, Diqun Yan |
IWDW | 4 |
| 2014 | Detecting Fake-Quality WAV Audio Based on Phase Differences
Jinglei Zhou, Rangding Wang, Chao Jin 0003, Diqun Yan |
IWDW | 4 |
| 2014 | A multipurpose audio aggregation watermarking based on multistage vector quantization
Rangding Wang, Diqun Yan, Youming Li |
Multim. Tools Appl. | 3 |
| 2014 | Detection of MP3Stego exploiting recompression calibration-based feature
Diqun Yan, Rangding Wang |
Multim. Tools Appl. | 1 |
| 2013 | A Huffman Table Index Based Approach to Detect Double MP3 Compression
Rangding Wang, Diqun Yan, Chao Jin 0003 |
IWDW | 3 |
| 2012 | Steganography for MP3 audio by exploiting the rule of window switching
Diqun Yan, Rangding Wang, Xianmin Yu |
Comput. Secur. | 1 |
| 2011 | Huffman table swapping-based steganograpy for MP3 audio
Diqun Yan, Rangding Wang |
Multim. Tools Appl. | 1 |
| 2009 | Quantization Step Parity-based Steganography for MP3 AudioabstractPetitcolas has proposed a steganographic technique called MP3Stego which can hide secret messages in a MP3 audio. This technique is well-known because of its high capacity. However, in rare cases, the normal audio encoding process will be terminated due to the endless loop problem caused by embedding operation. In addition, the statistical undetectability of MP3Stego can be further improved. Inspired by MP3Stego, a new steganographic method for MP3 audio is proposed in this paper. The parity bit of quantization step rather than the parity bit of block size in MP3Stego is employed to embed secret messages. Compared with MP3Stego, the proposed method can avoid the endless loop problem and achieve better imperceptibility and higher security. Diqun Yan, Rangding Wang, Liguang Zhang |
Fundam. Informaticae | 1 |