Qing Wang 0039

dblp:97/6505-39 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
13since 2021 · last 2025
0009-0008-5449-4815ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2025 DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
abstract
Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for humans and machines to detect. In this study, we propose DiffAttack, a novel timbre-reserved adversarial attack approach, that exploits the capability of a diffusion-based voice conversion (DiffVC) model to generate adversarial fake audio with distinct target speaker attribution. By introducing adversarial constraints into the diffusion-based voice conversion model’s generative process, we aim to craft fake samples that effectively mislead target models while preserving the speaker-wised characteristics. Specifically, inspired by the utilization of randomly sampled Gaussian noise in conventional adversarial attack and diffusion processes, we incorporate adversarial constraints into the reverse diffusion process. As a result, these adversarial constraints subtly guide the reverse diffusion process toward aligning with the target speaker distribution. Our experiments on the LibriTTS dataset indicate that our proposed DiffAttack significantly improves the attack success rate compared to vanilla DiffVC or other methods. Furthermore, objective and subjective evaluations demonstrate that introducing adversarial constraints does not compromise the speech quality generated by the DiffVC model.
Qing Wang 0039, Jixun Yao, Zhaokai Sun, Lei Xie 0001, John H. L. Hansen
ICASSP1
2025 Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
Zhaokai Sun, Qing Wang 0039, Lei Xie 0001
INTERSPEECH3
2025 Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
Hongjie Chen 0001, Qing Wang 0039, Hang Lv 0006, Jian Kang 0006, Jie Li 0001, Zhennan Lin, Lei Xie 0001
INTERSPEECH3
2024 Whisper-SV: Adapting Whisper for low-data-resource speaker verification
Li Zhang 0106, Qing Wang 0039, Lei Xie 0001
Speech Commun.3
2024 Distinctive and Natural Speaker Anonymization via Singular Value Transformation-Assisted Matrix
abstract
Speaker anonymization is an effective privacy protection solution that aims to conceal speaker's identity while preserving the naturalness and distinctiveness of the original speech. Mainstream approaches use an utterance-level vector from a pre-trained automatic speaker verification (ASV) model to represent speaker identity, which is then averaged or modified for anonymization. However, these systems suffer from deterioration in thenaturalnessof anonymized speech, degradation in speakerdistinctiveness, and severe privacy leakage against powerful attackers. To address these issues and especially generate more natural and distinctive anonymized speech, we propose a novel speaker anonymization approach that models a matrix related to speaker identity and transforms it into an anonymized singular value transformation-assisted matrix to conceal the original speaker identity. Our approach extracts frame-level speaker vectors from a pre-trained ASV model and employs an attention mechanism to create a speaker-score matrix and speaker-related tokens. Notably, the speaker-score matrix acts as the weight for the corresponding speaker-related token, representing the speaker's identity. The singular value transformation-assisted matrix is generated through the recomposition of the decomposed orthonormal eigenvectors matrix and non-linear transformed singular through Singular Value Decomposition (SVD). This process prevents the degradation of speaker distinctiveness caused by the introduction of other speakers' identity information. By multiplying the singular value transformation-assisted matrix and speaker-related tokens, we generate the anonymized speaker identity representation, thereby producing anonymized speech that is both natural and distinctive. Experiments on VoicePrivacy Challenge datasets demonstrate the effectiveness of our approach in protecting speaker privacy under all attack scenarios while maintaining speech naturalness and distinctiveness.
Jixun Yao, Qing Wang 0039, Ziqian Ning, Lei Xie 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Preserving Background Sound in Noise-Robust Voice Conversion Via Multi-Task Learning
abstract
Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare attention to VC with background sound. The critical problem for preserving background sound in VC is inevitable speech distortion by the neural separation model and the cascade mismatch between the source separation model and the VC model. In this paper, we propose an end-to-end framework via multitask learning which sequentially cascades a source separation (SS) module, a bottleneck feature extraction module and a VC module. Specifically, the source separation task explicitly considers critical phase information and limits the distortion caused by the imperfect separation process. The source separation task, the typical VC task and the unified task share a uniform reconstruction loss constrained by joint training to reduce the mismatch between the SS and VC modules. Experimental results demonstrate that our proposed framework significantly outperforms the baseline systems while achieving comparable quality and speaker similarity to the VC models trained with clean data.
Jixun Yao, Qing Wang 0039, Ziqian Ning, Lei Xie 0001, Danming Xie
ICASSP3
2023 Distinguishable Speaker Anonymization Based on Formant and Fundamental Frequency Scaling
abstract
Speech data on the Internet are proliferating exponentially because of the emergence of social media, and the sharing of such personal data raises obvious security and privacy concerns. One solution to mitigate these concerns involves concealing speaker identities before sharing speech data, also referred to as speaker anonymization. In our previous work, we have developed an automatic speaker verification (ASV)-model-free anonymization framework to protect speaker privacy while preserving speech intelligibility. Although the framework ranked first place in VoicePrivacy 2022 challenge, the anonymization was imperfect, since the speaker distinguishability of the anonymized speech was deteriorated. To address this issue, in this paper, we directly model the formant distribution and fundamental frequency (F0) to represent speaker identity and anonymize the source speech by the uniformly scaling formant and F0. By directly scaling the formant and F0, the speaker distinguishability degradation of the anonymized speech caused by the introduction of other speakers is prevented. The experimental results demonstrate that our proposed framework can improve the speaker distinguishability and significantly outperforms our previous framework in voice distinctiveness. Furthermore, our proposed method can trade off the privacy-utility by using different scaling factors.
Jixun Yao, Qing Wang 0039, Lei Xie 0001, Namin Wang, Jie Liu 0097
ICASSP2
2023 Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification
abstract
The scarcity of labeled far-field speech is a constraint for training superior far-field speaker verification systems. In general, fine-tuning the model pre-trained on large-scale near- field speech through a small amount of far-field speech substantially outperforms training from scratch. However, the vanilla fine-tuning suffers from two limitations – catastrophic forgetting and overfitting. In this paper, we propose a weight transfer regularization (WTR) loss to constrain the distance of the weights between the pre-trained model and the fine-tuned model. With the WTR loss, the fine-tuning process takes advantage of the previously acquired discriminative ability from the large-scale near-field speech and avoids catastrophic for- getting. Meanwhile, the analysis based on the PAC-Bayes generalization theory indicates that the WTR loss makes the fine-tuned model have a tighter generalization bound, thus mitigating the overfitting problem. Moreover, three different norm distances for weight transfer are explored, which are L1-norm distance, L2-norm distance, and Max-norm distance. We evaluate the effectiveness of the WTR loss on VoxCeleb (pre-trained) and FFSVC (fine-tuned) datasets. Experimental results show that the distance-based weight transfer fine-tuning strategy significantly outperforms vanilla fine- tuning and other competitive domain adaptation methods.
Li Zhang 0084, Qing Wang 0039, Wei Rao 0002, Yannan Wang, Lei Xie 0001
ICASSP2
2023 Pseudo-Siamese Network based Timbre-reserved Black-box Adversarial Attack in Speaker Identification
Qing Wang 0039, Jixun Yao, Lei Xie 0001
INTERSPEECH1
2023 Timbre-Reserved Adversarial Attack in Speaker Identification
abstract
As a type of biometric identification, a speaker identification (SID) system is confronted with various kinds of attacks. The spoofing attacks typically imitate the timbre of the target speakers, while the adversarial attacks confuse the SID system by adding a well-designed adversarial perturbation to an arbitrary speech. Although the spoofing attack copies a similar timbre as the victim, it does not exploit the vulnerability of the SID model and may not make the SID system give the attacker's desired decision. As for the adversarial attack, despite the SID system can be led to a designated decision, it cannot meet the specified text or speaker timbre requirements for the specific attack scenarios. In this study, to make the attack in SID not only leverage the vulnerability of the SID model but also reserve the timbre of the target speaker, we propose a timbre-reserved adversarial attack in the speaker identification. We generate the timbre-reserved adversarial audios by adding an adversarial constraint during the different training stages of the voice conversion (VC) model. Specifically, the adversarial constraint is using the target speaker label to optimize the adversarial perturbation added to the VC model representations and is implemented by a speaker classifier joining in the VC model training. The adversarial constraint can help to control the VC model to generate the speaker-wised audio. Eventually, the inference of the VC model is the ideal adversarial fake audio, which is timbre-reserved and can fool the SID system. Experimental results on the Audio deepfake detection (ADD) challenge dataset indicate that our proposed method improves the attack success rate significantly compare with the vanilla VC model without additionally introducing an adversarial noise to the attack speech. Objective and subjective evaluations illustrate that the quality of fake audio generated by our proposed method is better than directly adding adversarial perturbation to the VC-generated audio. Furthermore, the analysis shows that our generated adversarial fake audios also meet the specified text and target speaker timbre-reserved requirements of the attacker.
Qing Wang 0039, Jixun Yao, Li Zhang 0106, Lei Xie 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Backend Ensemble for Speaker Verification and Spoofing Countermeasure
abstract
This paper describes the NPU system submitted to Spoofing Aware Speaker Verification Challenge 2022.We particularly focus on the backend ensemble for speaker verification and spoofing countermeasure from three aspects.Firstly, besides simple concatenation, we propose circulant matrix transformation and stacking for speaker embeddings and countermeasure embeddings.With the stacking operation of newly-defined circulant embeddings, we almost explore all the possible interactions between speaker embeddings and countermeasure embeddings.Secondly, we attempt different convolution neural networks to selectively fuse the embeddings' salient regions into channels with convolution kernels.Finally, we design parallel attention in 1D convolution neural networks to learn the global correlation in channel dimensions as well as to learn the important parts in feature dimensions.Meanwhile, we embed squeeze-and-excitation attention in 2D convolutional neural networks to learn the global dependence among speaker embeddings and countermeasure embeddings.Experimental results demonstrate that all the above methods are effective.After fusion of four well-trained models enhanced by the mentioned methods, the best SASV-EER, SPF-EER and SV-EER we achieve are 0.559%, 0.354% and 0.857% on the evaluation set respectively.Together with the above contributions, our submission system achieves the fifth place in this challenge.
Li Zhang 0084, Qing Wang 0039, Lei Xie 0001
INTERSPEECH4
2021 Duality Temporal-Channel-Frequency Attention Enhanced Speaker Representation Learning
abstract
The use of channel-wise attention in CNN based speaker representation networks has achieved remarkable performance in speaker verification (SV). But these approaches do simple averaging on time and frequency feature maps before channel-wise attention learning and ignore the essential mutual interaction among temporal, channel as well as frequency scales. To address this problem, we propose the Duality Temporal-Channel-Frequency (DTCF) attention to re-calibrate the channel-wise features with aggregation of global context on temporal and frequency dimensions. Specifically, the duality attention - time-channel (T-C) attention as well as frequency-channel (F-C) attention - aims to focus on salient regions along the T-C and F-C feature maps that may have more considerable impact on the global context, leading to more discriminative speaker representations. We evaluate the effectiveness of the proposed DTCF attention on the CN-Celeb and VoxCeleb datasets. On the CN-Celeb evaluation set, the EER/minDCF of ResNet34-DTCF are reduced by 0.63%/0.0718 compared with those of ResNet34-SE. On VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1-H evaluation sets, the EER/minDCF of ResNet34-DTCF achieve 0.36%/0.0263, 0.39%/0.0382 and 0.74%/0.0753 reductions compared with those of ResNet34-SE.
Li Zhang 0084, Qing Wang 0039, Lei Xie 0001
ASRU2
2021 Multi-Level Transfer Learning from Near-Field to Far-Field Speaker Verification
abstract
In far-field speaker verification, the performance of speaker embeddings is susceptible to degradation when there is a mismatch between the conditions of enrollment and test speech.To solve this problem, we propose the feature-level and instancelevel transfer learning in the teacher-student framework to learn a domain-invariant embedding space.For the feature-level knowledge transfer, we develop the contrastive loss to transfer knowledge from teacher model to student model, which can not only decrease the intra-class distance, but also enlarge the inter-class distance.Moreover, we propose the instance-level pairwise distance transfer method to force the student model to preserve pairwise instances distance from the well optimized embedding space of the teacher model.On FFSVC 2020 evaluation set, our EER on Full-eval trials is relatively reduced by 13.9% compared with the fusion system result on Partialeval trials of Task2.On Task1, compared with the winner's DenseNet result on Partial-eval trials, our minDCF on Full-eval trials is relatively reduced by 6.3%.On Task3, the EER and minDCF of our proposed method on Full-eval trials are very close to the result of the fusion system on Partial-eval trials.Our results also outperform other competitive domain adaptation methods.
Li Zhang 0084, Qing Wang 0039, Kong-Aik Lee, Lei Xie 0001, Haizhou Li 0001
Interspeech2
2020 Inaudible Adversarial Perturbations for Targeted Attack in Speaker Recognition
abstract
Speaker recognition is a popular topic in biometric authentication and many deep learning approaches have achieved extraordinary performances.However, it has been shown in both image and speech applications that deep neural networks are vulnerable to adversarial examples.In this study, we aim to exploit this weakness to perform targeted adversarial attacks against the x-vector based speaker recognition system.We propose to generate inaudible adversarial perturbations achieving targeted white-box attacks to speaker recognition system based on the psychoacoustic principle of frequency masking.Specifically, we constrict the perturbation under the masking threshold of original audio, instead of using a common lp norm to measure the perturbations.Experiments on Aishell-1 corpus show that our approach yields up to 98.5% attack success rate to arbitrary gender speaker targets, while retaining indistinguishable attribute to listeners.Furthermore, we also achieve an effective speaker attack when applying the proposed approach to a completely irrelevant waveform, such as music.
Qing Wang 0039, Lei Xie 0001
INTERSPEECH1
2019 I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared Experiences
abstract
The I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation.
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco
INTERSPEECH16
2019 Adversarial Regularization for End-to-End Robust Speaker Verification
Qing Wang 0039, Sining Sun, Lei Xie 0001, John H. L. Hansen
INTERSPEECH1
2018 Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker Recognition
abstract
The i-vector approach to speaker recognition has achieved good performance when the domain of the evaluation dataset is similar to that of the training dataset. However, in realworld applications, there is always a mismatch between the training and evaluation datasets, that leads to performance degradation. To address this problem, this paper proposes to learn the domain-invariant and speaker-discriminative speech representations via domain adversarial training. Specifically, with domain adversarial training method, we use a gradient reversal layer to remove the domain variation and project the different domain data into the same subspace. Moreover, we compare the proposed method with other state-of-the-art unsupervised domain adaptation techniques for i-vector approach to speaker recognition (e.g. autoencoder based domain adaptation, inter dataset variability compensation, dataset-invariant covariance normalization, and so on). Experiments on 2013 domain adaptation challenge (DAC) dataset demonstrate that the proposed method is not only effective in solving the dataset mismatch problem, but also outperforms the compared unsupervised domain adaptation methods.
Qing Wang 0039, Wei Rao 0002, Sining Sun, Lei Xie 0001, Chng Eng Siong, Haizhou Li 0001
ICASSP1