Chenglong Wang 0001

dblp:94/9817-1 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
17since 2021 · last 2025
0000-0002-5785-7027ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021
YearPublicationVenuePosition
2025 Region-Based Optimization in Continual Learning for Audio Deepfake Detection
abstract
Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real-world deepfakes. To address this issue, we propose a continual learning method named Region-Based Optimization (RegO) for audio deepfake detection. Specifically, we use the Fisher information matrix to measure important neuron regions for real and fake audio detection, dividing them into four regions. First, we directly fine-tune the less important regions to quickly adapt to new tasks. Next, we apply gradient optimization in parallel for regions important only to real audio detection, and in orthogonal directions for regions important only to fake audio detection. For regions that are important to both, we use sample proportion-based adaptive gradient optimization. This region-adaptive optimization ensures an appropriate trade-off between memory stability and learning plasticity. Additionally, to address the increase of redundant neurons from old tasks, we further introduce the Ebbinghaus forgetting mechanism to release them, thereby promoting the model’s ability to learn more generalized discriminative features. Experimental results show our method achieves a 21.3 percent improvement in EER over the state-of-the-art continual learning approach RWM for audio deepfake detection. Moreover, the effectiveness of RegO extends beyond the audio deepfake detection domain, showing potential significance in other tasks, such as image recognition.
Yujie Chen 0006, Jiangyan Yi, Cunhang Fan, Jianhua Tao 0001, Yong Ren 0006, Siding Zeng, Chu Yuan Zhang, Xinrui Yan, Jun Xue 0001, Chenglong Wang 0001, Zhao Lv, Xiaohui Zhang 0006
AAAI11
2025 ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
abstract
Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness. To this end, we propose ALLM4ADD, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: ''Is this audio fake or real?''. We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems. Code is available at https://github.com/ucas-hao/qwen_audio_for_add.git.
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Zheng Lian 0004, Yong Ren 0006, Yujie Chen 0006, Zhengqi Wen
ACM Multimedia3
2024 What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection
abstract
The rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms. Existing detection models have shown remarkable success in discriminating known deepfake audio, but struggle when encountering new attack types. To address this challenge, one of the emergent effective approaches is continual learning. In this paper, we propose a continual learning approach called Radian Weight Modification (RWM) for audio deepfake detection. The fundamental concept underlying RWM involves categorizing all classes into two groups: those with compact feature distributions across tasks, such as genuine audio, and those with more spread-out distributions, like various types of fake audio. These distinctions are quantified by means of the in-class cosine distance, which subsequently serves as the basis for RWM to introduce a trainable gradient modification direction for distinct data types. Experimental evaluations against mainstream continual learning methods reveal the superiority of RWM in terms of knowledge acquisition and mitigating forgetting in audio deepfake detection. Furthermore, RWM's applicability extends beyond audio deepfake detection, demonstrating its potential significance in diverse machine learning domains such as image recognition.
Xiaohui Zhang 0006, Jiangyan Yi, Chenglong Wang 0001, Chu Yuan Zhang, Siding Zeng, Jianhua Tao 0001
AAAI3
2024 Multi-Scale Permutation Entropy for Audio Deepfake Detection
abstract
With the widespread application of Automatic Speaker Verification (ASV) technology in security authentication, the threat of fake audio attacks looms as a malicious means compromising system security. In this study, we employ the multi-scale permutation entropy (MPE) in audio deepfake detection, which could help measure the complexity and detect the dynamic characteristics of audio signals at different scales. Experimental results indicate that MPE can effectively improve the performance of LFCC. For example, on the ASVspoof2019 LA test set, it successfully achieves an equal error rate (EER) of less than 2%, which is around 50% lower than that of LFCC. Notably, MPE exhibits extraordinary generalization performance when applied to the In-the-Wild dataset, as its performance of EER is comparable to that of Wav2vec, without requiring pretraining. Therefore, we believe that MPE holds promising prospects in voice biometric recognition for anti-spoofing applications. Our code is available at https://github.com/ADDchallenge/MPE-for-audio-deepfake-detection
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Xiaohui Zhang 0006
ICASSP1
2024 RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
abstract
Fake artefacts for discriminating between bonafide and fake audio can exist in both short-and long-range segments.Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio.This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short-and long-range discriminative information for audio deepfake detection.Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information.Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining shortand long-range information.The results show that our proposed RawBMamba achieves a 34.1% improvement over Rawformer on ATSVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.Codes will be released on https://github.com/cyjie429/RawBMamba.
Yujie Chen 0006, Jiangyan Yi, Jun Xue 0001, Chenglong Wang 0001, Xiaohui Zhang 0006, Shunbo Dong, Siding Zeng, Jianhua Tao 0001, Zhao Lv, Cunhang Fan
INTERSPEECH4
2024 Utilizing Speaker Profiles for Impersonation Audio Detection
abstract
Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice conversion (VC). However, countermeasures against impersonation remain an underexplored area. Impersonation is a fake type that involves an imitator replicating specific traits and speech style of a target speaker. Unlike TTS and VC, which often leave digital traces or signal artifacts, impersonation involves live human beings producing entirely natural speech, rendering the detection of impersonation audio a challenging task. Thus, we propose a novel method that integrates speaker profiles into the process of impersonation audio detection. Speaker profiles are inherent characteristics that are challenging for impersonators to mimic accurately, such as speaker's age, job. We aim to leverage these features to extract discriminative information for detecting impersonation audio. Moreover, there is no large impersonated speech corpora available for quantitative study of impersonation impacts. To address this gap, we further design the first large-scale, diverse-speaker Chinese impersonation dataset, named ImPersonation Audio Detection (IPAD), to advance the community's research on impersonation audio detection. We evaluate several existing fake audio detection methods on our proposed dataset IPAD, demonstrating its necessity and the challenges. Additionally, our findings reveal that incorporating speaker profiles can significantly enhance the model's performance in detecting impersonation audio.
Jiangyan Yi, Chenglong Wang 0001, Yong Ren 0006, Jianhua Tao 0001, Xinrui Yan, Yujie Chen 0006, Xiaohui Zhang 0006
ACM Multimedia3
2024 Spatial reconstructed local attention Res2Net with F0 subband for fake speech detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Chenglong Wang 0001, Chengshi Zheng, Zhao Lv
Neural Networks5
2024 SceneFake: An initial dataset and benchmarks for scene fake audio detection
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Chuyuan Zhang, Cunhang Fan, Zhengkun Tian, Haoxin Ma, Ruibo Fu
Pattern Recognit.2
2024 CFAD: A Chinese dataset for fake audio detection
Haoxin Ma, Jiangyan Yi, Chenglong Wang 0001, Xinrui Yan, Jianhua Tao 0001, Tao Wang 0074, Ruibo Fu
Speech Commun.3
2023 Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
abstract
In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on, which are often perceived by shallow networks. However, shallow networks have much noise, which can not capture this very well. To address this problem, we propose using the deepest network instruct shallow network for enhancing shallow networks. Specifically, the networks of FSD are divided into several segments, the deepest network being used as the teacher model, and all shallow networks become multiple student models by adding classifiers. Meanwhile, the distillation path between the deepest network feature and shallow network features is used to reduce the feature difference. A series of experimental results on the ASVspoof 2019 LA and PA datasets show the effectiveness of the proposed method, with significant improvements compared to the baseline.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Chenglong Wang 0001, Zhengqi Wen, Dan Zhang 0014, Zhao Lv
ICASSP4
2023 Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection
abstract
Current fake audio detection algorithms have achieved promising performances on most datasets. However, their performance may be significantly degraded when dealing with audio of a different dataset. The orthogonal weight modification to overcome catastrophic forgetting does not consider the similarity of genuine audio across different datasets. To overcome this limitation, we propose a continual learning algorithm for fake audio detection to overcome catastrophic forgetting, called Regularized Adaptive Weight Modification (RAWM). When fine-tuning a detection network, our approach adaptively computes the direction of weight modification according to the ratio of genuine utterances and fake utterances. The adaptive modification direction ensures the network can effectively detect fake audio on the new dataset while preserving its knowledge of old model, thus mitigating catastrophic forgetting. In addition, genuine audio collected from quite different acoustic conditions may skew their feature distribution, so we introduce a regularization constraint to force the network to remember the old distribution in this regard. Our method can easily be generalized to related fields, like speech emotion recognition. We also evaluate our approach across multiple datasets and obtain a significant performance improvement on cross-dataset experiments.
Xiaohui Zhang 0006, Jiangyan Yi, Jianhua Tao 0001, Chenglong Wang 0001, Chu Yuan Zhang
ICML4
2023 Detection of Cross-Dataset Fake Audio Based on Prosodic and Pronunciation Features
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Shuai Zhang 0014, Xun Chen 0001
INTERSPEECH1
2023 TO-Rawnet: Improving RawNet with TCN and Orthogonal Regularization for Fake Audio Detection
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Shuai Zhang 0014, Ruibo Fu, Xun Chen 0001
INTERSPEECH1
2023 Adversarial Multi-Task Learning for Mandarin Prosodic Boundary Prediction With Multi-Modal Embeddings
abstract
Prosodic boundaries are still crucial to the naturalness of end-to-end speech synthesis systems. This article proposes to use adversarial multi-task learning to predict prosodic boundaries. Adversarial multi-task learning is utilized to transfer knowledge from an auxiliary POS tagging task to a prosodic boundary prediction task. Furthermore, multi-modal embeddings are composed of contextual word and speech embedding features obtained from the pre-trained bidirectional encoder representations from transformers (BERT) model and Speech2Vec. We can utilize linguistic and acoustic information from large amounts of external text and speech data without prosodic boundary labels. At the inference stage, the prosodic boundary predicting model can use the syntactic features learnt from the POS tagging task without any extra computation cost due to only employing the prosodic boundary predicting task to decode. We conducted experiments on Mandarin datasets. The results show that the models using multi-modal embeddings from the pre-trained BERT and Speech2Vec outperform the models trained with single modal embedding. Furthermore, the models trained with adversarial training obtain further performance gains by up to 2.95% in$F_{1}$score.
Jiangyan Yi, Jianhua Tao 0001, Ruibo Fu, Tao Wang 0074, Chu Yuan Zhang, Chenglong Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2022 ADD 2022: the first Audio Deep Synthesis Detection Challenge
abstract
Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks.
Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001
ICASSP6
2021 Continual Learning for Fake Audio Detection
abstract
Fake audio attack becomes a major threat to the speaker verification system. Although current detection approaches have achieved promising results on dataset-specific scenarios, they encounter difficulties on unseen spoofing data. Fine-tuning and retraining from scratch have been applied to incorporate new data. However, fine-tuning leads to performance degradation on previous data. Retraining takes a lot of time and computation resources. Besides, previous data are unavailable due to privacy in some situations. To solve the above problems, this paper proposes detecting fake without forgetting, a continual-learning-based method, to make the model learn new spoofing attacks incrementally. A knowledge distillation loss is introduced to loss function to preserve the memory of original model. Supposing the distribution of genuine voice is consistent among different scenarios, an extra embedding similarity loss is used as another constraint to further do a positive sample alignment. Experiments are conducted on the ASVspoof2019 dataset. The results show that our proposed method outperforms fine-tuning by the relative reduction of average equal error rate up to 81.62%.
Haoxin Ma, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Zhengkun Tian, Chenglong Wang 0001
Interspeech6
2021 Half-Truth: A Partially Fake Audio Detection Dataset
abstract
Diverse promising datasets have been designed to further the development of fake audio detection, such as ASVspoof databases.However, previous datasets ignore an attacking situation, in which the hacker hides some small fake clips in real speech audio.This poses a serious threat since that it is difficult to distinguish the small fake clip from the whole speech utterance.Therefore, this paper develops such a dataset for half-truth audio detection (HAD).Partially fake audio in the HAD dataset involves only changing a few words in an utterance.The audio of the words is generated with the very latest state-of-the-art speech synthesis technology.We can not only detect fake uttrances but also localize manipulated regions in a speech using this dataset.Some benchmark results are presented on this dataset.The results show that partially fake audio presents much more challenging than fully fake audio for fake audio detection.The HAD dataset is publicly available 1 .
Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Haoxin Ma, Zhengkun Tian, Chenglong Wang 0001, Tao Wang 0074, Ruibo Fu
Interspeech6