Na Li 0012

dblp:18/3173-12 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator
abstract
Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges, manifesting as mispronunciations, audible noise, and quality degradation. To address these issues, we introduce Vox-Evaluator, a multi-level evaluator designed to guide the correction of erroneous speech segments and preference alignment for TTS systems. It is capable of identifying the temporal boundaries of erroneous segments and providing a holistic quality assessment of the generated speech. Specifically, to refine erroneous segments and enhance the robustness of the zero-shot TTS model, we propose to automatically identify acoustic errors with the evaluator, mask the erroneous segments, and finally regenerate speech conditioning on the correct portions. In addition, the fine-gained information obtained from Vox-Evaluator can guide the preference alignment for TTS model, thereby reducing the bad cases in speech synthesize. Due to the lack of suitable training datasets for the Vox-Evaluator, we also constructed a synthesized text-speech dataset annotated with fine-grained pronunciation errors or audio quality issues. The experimental results demonstrate the effectiveness of the proposed Vox-Evaluator in enhancing the stability and fidelity of TTS systems through the speech correction mechanism and preference optimization.
Hualei Wang, Na Li 0012, Chuke Wang, Zhifeng Li 0001, Dong Yu 0001
AAAI2
2024 Investigating Long-Term and Short-Term Time-Varying Speaker Verification
abstract
The performance of speaker verification systems can be adversely affected by time domain variations. However, limited research has been conducted on time-varying speaker verification due to the absence of appropriate datasets. This paper aims to investigate the impact of long-term and short-term time-varying in speaker verification and proposes solutions to mitigate these effects. For long-term speaker verification (i.e., cross-age speaker verification), we introduce an age-decoupling adversarial learning method to learn age-invariant speaker representation by mining age information from the VoxCeleb dataset. For short-term speaker verification, we collect the SMIIP-TimeVarying (SMIIP-TV) Dataset, which includes recordings at multiple time slots every day from 373 speakers for 90 consecutive days and other relevant meta information. Using this dataset, we analyze the time-varying of speaker embeddings and propose a novel but realistic time-varying speaker verification task, termed incremental sequence-pair speaker verification. This task involves continuous interaction between enrollment audios and a sequence of testing audios with the aim of improving performance over time. We introduce the template updating method to counter the negative effects over time, and then formulate the template updating processing as a Markov Decision Process and propose a template updating method based on deep reinforcement learning (DRL). The policy network of DRL is treated as an agent to determine if and how much should the template be updated. In summary, this paper releases our collected database, investigates both the long-term and short-term time-varying scenarios and provides insights and solutions into time-varying speaker verification.
Xiaoyi Qin, Na Li 0012, Shufei Duan, Ming Li 0026
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Simple Attention Module Based Speaker Verification with Iterative Noisy Label Detection
abstract
Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (SimAM), for speaker verification. The SimAM module is a plug-and-play module without extra modal parameters. In addition, we propose a noisy label detection method to iteratively filter out the data samples with a noisy label from the training data, considering that a large-scale dataset labeled with human annotation or other automated processes may contain noisy labels. Data with the noisy label may over parameterize a deep neural network (DNN) and result in a performance drop due to the memorization effect of the DNN. Experiments are conducted on VoxCeleb dataset. The speaker verification model with SimAM achieves the 0.675% equal error rate (EER) on VoxCeleb1 original test trials. Our proposed iterative noisy label detection method further reduces the EER to 0.643%.
Xiaoyi Qin, Na Li 0012, Chao Weng, Dan Su 0002, Ming Li 0026
ICASSP2
2022 The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
abstract
This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech recognition (ASR) tasks. In these meeting scenarios, the uncertainty of the speaker number and the high ratio of overlapped speech present great challenges for diarization. Based on the assumption that there is valuable complementary information between acoustic features, spatial-related and speaker-related features, we propose a multi-level feature fusion mechanism based target-speaker voice activity detection (FFM-TS-VAD) system to improve the performance of the conventional TS-VAD system. Furthermore, we propose a data augmentation method during training to improve the system robustness when the angular difference between two speakers is relatively small. We provide comparisons for different sub-systems we used in M2MeT challenge. Our submission is a fusion of several sub-systems and ranks second in the diarization task.
Naijun Zheng, Na Li 0012, Xixin Wu, Lingwei Meng, Jiawen Kang 0002, Chao Weng, Dan Su 0002, Helen M. Meng
ICASSP2
2022 Multi-Channel Speaker Diarization Using Spatial Features for Meetings
abstract
Speaker identification for overlapped speech presents a great challenge for speaker diarization tasks in meeting scenarios. In order to overcome such challenges, several overlap-aware resegmentation methods based on deep learning have been integrated into speaker diarization systems. In this paper we propose two multi-channel diarization systems which have enhanced capability in detecting overlapped speech and identify speakers via learning spatial features. The first system applies a multi-look strategy to train networks without given the speakers’ direction of arrival(DOA), and the other system estimates the DOA of target speakers based on existing diarization results. Both systems aim to estimate the voice activity of speakers in different directions to handle overlapped speech. Experimental results on the AMI corpus show that the relative improvements of both systems can reach 9.4% and 18.1% in term of diarization error rate (DER) against an overlap-aware single-channel system with a BeamformIt front-end.
Naijun Zheng, Na Li 0012, Jianwei Yu 0001, Chao Weng, Dan Su 0002, Xunying Liu, Helen M. Meng
ICASSP2
2022 Cross-Age Speaker Verification: Learning Age-Invariant Speaker Embeddings
abstract
Automatic speaker verification has achieved remarkable progress in recent years.However, there is little research on cross-age speaker verification (CASV) due to insufficient relevant data.In this paper, we mine cross-age test sets based on the VoxCeleb dataset and propose our age-invariant speaker representation(AISR) learning method.Since the VoxCeleb is collected from the YouTube platform, the dataset consists of crossage data inherently.However, the meta-data does not contain the speaker age label.Therefore, we adopt the face age estimation method to predict the speaker age value from the associated visual data, then label the audio recording with the estimated age.We construct multiple Cross-Age test sets on VoxCeleb (Vox-CA), which deliberately select the positive trials with large age-gap.Also, the effect of nationality and gender is considered in selecting negative pairs to align with Vox-H cases.The baseline system performance drops from 1.939% EER on the Vox-H test set to 10.419% on the Vox-CA20 test set, which indicates how difficult the cross-age scenario is.Consequently, we propose an age-decoupling adversarial learning (ADAL) method to alleviate the negative effect of the age gap and reduce intra-class variance.Our method outperforms the baseline system by over 10% related EER reduction on the Vox-CA20 test set.The source code and trial resources are available on https://github.com/qinxiaoyi/Cross-AgeSpeaker Verification.
Xiaoyi Qin, Na Li 0012, Chao Weng, Dan Su 0002, Ming Li 0026
INTERSPEECH2
2021 Replay and Synthetic Speech Detection with Res2Net Architecture
abstract
Existing approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-called Res2Net, to improve the anti-spoofing countermeasure’s generalizability. Res2Net mainly modifies the ResNet block to enable multiple feature scales. Specifically, it splits the feature maps within one block into multiple channel groups and designs a residual-like connection across different channel groups. Such connection increases the possible receptive fields, resulting in multiple feature scales. This multiple scaling mechanism significantly improves the countermeasure’s generalizability to unseen spoofing attacks. It also decreases the model size compared to ResNet-based models. Experimental results show that the Res2Net model consistently outperforms ResNet34 and ResNet50 by a large margin in both physical access (PA) and logical access (LA) of the ASVspoof 2019 corpus. Moreover, integration with the squeeze-and-excitation (SE) block can further enhance performance. For feature engineering, we investigate the gen-eralizability of Res2Net combined with different acoustic features, and observe that the constant-Q transform (CQT) achieves the most promising performance in both PA and LA scenarios. Our best single system outperforms other state-of-the-art single systems in both PA and LA of the ASVspoof 2019 corpus.
Xu Li 0015, Na Li 0012, Chao Weng, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
ICASSP2
2021 A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech
abstract
In multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint training framework of the front-end multi-look speech separator and the back-end speaker embedding extractor is proposed for multi-channel overlapped speech. To better leverage the complementarity between the speech separator and the speaker embedding extractor, several training strategies are proposed to jointly optimize the two modules. Experimental results show that the proposed joint training framework significantly outperforms the individual SV system by around 52% relative EER reduction. Additionally, the robustness of the proposed framework is further evaluated under different conditions.
Naijun Zheng, Na Li 0012, Bo Wu 0011, Meng Yu 0003, Jianwei Yu 0001, Chao Weng, Dan Su 0002, Xunying Liu, Helen M. Meng
ICASSP2
2020 Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization
abstract
Adapting speaker verification (SV) systems to a new environment is a very challenging task. Current adaptation methods in SV mainly focus on the backend, i.e, adaptation is carried out after the speaker embeddings have been created. In this paper, we present a DNN-based adaptation method using maximum mean discrepancy (MMD). Our method exploits two important aspects neglected by previous research. First, instead of minimizing domain discrepancy at utterance-level alone, our method minimizes domain discrepancy at both frame-level and utterance-level, which we believe will make the adaptation more robust to the duration discrepancy between training data and test data. Second, we introduce a consistency regularization for unlabelled target-domain data. The consistency regularization encourages the target speaker embeddings robust to adverse perturbations. Experiments on NIST SRE 2016 and 2018 show that our DNN adaptation works significantly better than the previously proposed DNN adaptation methods. What's more, our method works well with backend adaptation. By combining the proposed method with backend adaptation, we achieve a 9% improvement over backend adaptation in SRE18.
Weiwei Lin 0002, Man-Wai Mak, Na Li 0012, Dan Su 0002, Dong Yu 0001
ICASSP3
2020 Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
abstract
Recently adversarial attacks on automatic speaker verification (ASV) systems attracted widespread attention as they pose severe threats to ASV systems.However, methods to defend against such attacks are limited.Existing approaches mainly focus on retraining ASV systems with adversarial data augmentation.Also, countermeasure robustness against different attack settings are insufficiently investigated.Orthogonal to prior approaches, this work proposes to defend ASV systems against adversarial attacks with a separate detection network, rather than augmenting adversarial data into ASV training.A VGG-like binary classification detector is introduced and demonstrated to be effective on detecting adversarial samples.To investigate detector robustness in a realistic defense scenario where unseen attack settings may exist, we analyze various kinds of unseen attack settings' impact and observe that the detector is robust (6.27%EER det degradation in the worst case) against unseen substitute ASV systems, but it has weak robustness (50.37%EER det degradation in the worst case) against unseen perturbation methods.The weak robustness against unseen perturbation methods shows a direction for developing stronger countermeasures.
Xu Li 0015, Na Li 0012, Jinghua Zhong, Xixin Wu, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH2
2020 A Framework for Adapting DNN Speaker Embedding Across Languages
abstract
Language mismatch remains a major hindrance to the extensive deployment of speaker verification (SV) systems. Current language adaptation methods in SV mainly rely on linear projection in embedding space; i.e., adaptation is carried out after the speaker embeddings have been created, which underutilizes the powerful representation of deep neural networks. This article proposes a maximum mean discrepancy (MMD) based framework for adapting deep neural network (DNN) speaker embedding across languages, featuring multi-level domain loss, separate batch normalization, and consistency regularization. We refer to the framework as MSC. We show that (1) minimizing domain discrepancy at both frame- and utterance-levels performs significantly better than at utterance-level alone; (2) separating the source-domain data from the target-domain in batch normalization improves adaptation performance; and (3) data augmentation can be utilized in the unlabelled target-domain through consistency regularization. By combining these findings, we achieve an EER of 8.69% and 7.95% in NIST SRE 2016 and 2018, respectively, which are significantly better than the previously proposed DNN adaptation methods. Our framework also works well with backend adaptation. By combining the proposed framework with backend adaptation, we achieve an 11.8% improvement over the backend adaptation in SRE18. When applying our framework to a 121-layer Densenet, we achieved an EER of 7.81% and 7.02% in NIST SRE 2016 and 2018, respectively.
Weiwei Lin 0002, Man-Wai Mak, Na Li 0012, Dan Su 0002, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Syllable-Dependent Discriminative Learning for Small Footprint Text-Dependent Speaker Verification
abstract
This study proposes a novel scheme of syllable-dependent discriminative speaker embedding learning for small footprint text-dependent speaker verification systems. To suppress undesired syllable variation and enhance the power of discrimination inherited in the frame-level features, we design a novel syllable-dependent clustering loss to optimize the network. Specifically, this loss function utilizes syllable labels as auxiliary supervision information to explicitly maximize inter-syllable divisibility and intra-syllable compactness between the learned frame-level features. Successively, we propose two syllable-dependent pooling mechanisms to aggregate the frame-level features to several syllable-level features by averaging those features corresponding to each syllable. The utterance-level speaker embeddings with powerful discrimination are then obtained by concatenating the syllable-level features. Experimental results on Tencent voice wake-up dataset show that our proposed scheme can accelerate the network convergence and achieve significant performance improvement against the state-of-the-art methods.
Junyi Peng, Yuexian Zou, Na Li 0012, Deyi Tuo, Dan Su 0002, Meng Yu 0003, Dong Yu 0001
ASRU3
2019 Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification
abstract
Deep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large Margin Cosine Loss (LMCL) in face recognition, we propose an enhanced LMCL named boundary discriminative LMCL (BD-LMCL) to emphasize the discriminative information inherited in the speaker boundaries. Unlike LMCL, where all training samples contribute equally for the objective function, only the samples around the speaker boundaries are considered during the network training with BD-LMCL. Specifically, those samples close to the boundaries are dynamically selected using top-k zero-one loss. Experimental results on a short duration corpus Android Cellphone and NIST SRE 2012 demonstrate better performance compared to LMCL and other popular loss functions.
Rongjin Li, Na Li 0012, Deyi Tuo, Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP2
2019 Learning Discriminative Features in Sequence Training without Requiring Framewise Labelled Data
abstract
In this work, we try to answer two questions: Can deeply learned features with discriminative power benefit an ASR system’s robustness to acoustic variability? And how to learn them without requiring framewise labelled sequence training data? As existing methods usually require knowing where the labels occur in the input sequence, they have so far been limited to many real-world sequence learning tasks. We propose a novel method which simultaneously models both the sequence discriminative training and the feature discriminative learning within a single network architecture, so that it can learn discriminative deep features in sequence training that obviates the need for presegmented training data. Our experiment in a realistic industrial ASR task shows that, without requiring any specific fine-tuning or additional complexity, our proposed models have consistently outperformed state-of-the-art models and significantly reduced Word Error Rate (WER) under all test conditions, and especially with highest improvements under unseen noise conditions, by relative 12.94%, 8.66% and 5.80%, showing our proposed models can generalize better to acoustic variability.
Jun Wang 0091, Dan Su 0002, Jie Chen 0057, Shulin Feng, Dongpeng Ma, Na Li 0012, Dong Yu 0001
ICASSP6
2019 Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification
abstract
In this paper, we present a Sequence-to-Sequence Attentional Siamese Neural Network (Seq2Seq-ASNN) that leverages temporal alignment information for end-to-end speaker verification. In prior works of speaker discriminative neural networks, utterance-level evaluation/enrollment speaker representations are usually calculated. Our proposed model, utilizing a sequence-to-sequence (Seq2Seq) attention mechanism, maps the frame-level evaluation representation into enrollment feature domain and further generates an utterance-level evaluation-enrollment joint vector for final similarity measure. Feature learning, attention mechanism, and metric learning are jointly optimized using an end-to-end loss function. Experimental results show that our proposed model outperforms various baseline methods, including the traditional i-Vector/PLDA method, multi-enrollment end-to-end speaker verification models, d-vector approaches, and a self attention model, for text-dependent speaker verification on a Tencent internal voice wake-up dataset.
Meng Yu 0003, Na Li 0012, Chengzhu Yu, Jia Cui, Dong Yu 0001
ICASSP3
2018 Deep Discriminative Embeddings for Duration Robust Speaker Verification
Na Li 0012, Deyi Tuo, Dan Su 0002, Zhifeng Li 0001, Dong Yu 0001
INTERSPEECH1
2017 Discriminative subspace modeling of SNR and duration variabilities for robust speaker verification
Na Li 0012, Man-Wai Mak, Weiwei Lin 0002, Jen-Tzung Chien
Comput. Speech Lang.1
2017 DNN-Driven Mixture of PLDA for Robust Speaker Verification
abstract
The mismatch between enrollment and test utterances due to different types of variabilities is a great challenge in speaker verification. Based on the observation that the SNR-level variability or channel-type variability causes heterogeneous clusters in i-vector space, this paper proposes to apply supervised learning to drive or guide the learning of probabilistic linear discriminant analysis (PLDA) mixture models. Specifically, a deep neural network (DNN) is trained to produce the posterior probabilities of different SNR levels or channel types given i-vectors as input. These posteriors then replace the posterior probabilities of indicator variables in the mixture of PLDA. The discriminative training causes the mixture model to perform more reasonable soft divisions of the i-vector space as compared to the conventional mixture of PLDA. During verification, given a test i-vector and a target-speaker's i-vector, the marginal likelihood for the same-speaker hypothesis is obtained by summing the component likelihoods weighted by the component posteriors produced by the DNN, and likewise for the different-speaker hypothesis. Results based on NIST 2012 SRE demonstrate that the proposed scheme leads to better performance under more realistic situations where both training and test utterances cover a wide range of SNRs and different channel types. Unlike the previous SNR-dependent mixture of PLDA which only focuses on SNR mismatch, the proposed model is more general and is potentially applicable to addressing different types of variability in speech.
Na Li 0012, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 SNR-invariant PLDA with multiple speaker subspaces
abstract
To deal with the mismatch between the enrollment and test utterances caused by noise with different signal-to-noise ratios (SNR), we have recently proposed an SNR-invariant PLDA model for robust speaker verification. In the model, SNR-specific information were separated from speaker-specific information through marginalizing out the SNR factors during the scoring process. However, this modeling approach assumes that speaker variabilities can be captured by a single speaker subspace regardless of the noise level of the utterances. We will show in this paper that i-vectors extracted from utterances with different noise levels will shift to different regions of the i-vector space and that i-vectors extracted from utterances having similar SNR tend to cluster together. In view of this observation, we propose introducing multiple speaker subspaces to the SNR-invariance PLDA model and use multiple covariance matrices to represent SNR-dependent channel variability. Through NIST 2012 SRE, this paper demonstrates that this finer and more precise modeling of speaker and SNR variabilities leads to better performance when compared with the conventional PLDA and SNR-invariant PLDA.
Na Li 0012, Man-Wai Mak
ICASSP1
2016 Deep neural network driven mixture of PLDA for robust i-vector speaker verification
abstract
In speaker recognition, the mismatch between the enrollment and test utterances due to noise with different signal-to-noise ratios (SNRs) is a great challenge. Based on the observation that noise-level variability causes the i-vectors to form heterogeneous clusters, this paper proposes using an SNR-aware deep neural network (DNN) to guide the training of PLDA mixture models. Specifically, given an i-vector, the SNR posterior probabilities produced by the DNN are used as the posteriors of indicator variables of the mixture model. As a result, the proposed model provides a more reasonable soft division of the i-vector space compared to the conventional mixture of PLDA. During verification, given a test trial, the marginal likelihoods from individual PLDA models are linearly combined by the posterior probabilities of SNR levels computed by the DNN. Experimental results for SNR mismatch tasks based on NIST 2012 SRE suggest that the proposed model is more effective than PLDA and conventional mixture of PLDA for handling heterogeneous corpora.
Na Li 0012, Man-Wai Mak, Jen-Tzung Chien
SLT1
2015 SNR-invariant PLDA modeling for robust speaker verification
abstract
16th Annual Conference of the International Speech Communication Association, INTERSPEECH 2015, Dresden, Germany, September 6-10, 2015
Na Li 0012, Man-Wai Mak
INTERSPEECH1
2015 SNR-Invariant PLDA Modeling in Nonparametric Subspace for Robust Speaker Verification
abstract
While i-vector/PLDA framework has achieved great success, its performance still degrades dramatically under noisy conditions. To compensate for the variability of i-vectors caused by different levels of background noise, this paper proposes an SNR-invariant PLDA framework for robust speaker verification. First, nonparametric feature analysis (NFA) is employed to suppress intra-speaker variation and emphasize the discriminative information inherited in the boundaries between speakers in the i-vector space. Then, in the NFA-projected subspace, SNR-invariant PLDA is applied to separate the SNR-specific information from speaker-specific information using an identity factor and an SNR factor. Accordingly, a projected i-vector in the NFA subspace can be represented as a linear combination of three components: speaker, SNR, and channel. During verification, the variability due to SNR and channels are integrated out when computing the marginal likelihood ratio. Experiments based on NIST 2012 SRE show that the proposed framework achieves superior performance when compared with the conventional PLDA and SNR-dependent mixture of PLDA.
Na Li 0012, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Clustering similar acoustic classes in the Fishervoice framework
abstract
In the Fishervoice (FSH) based framework, the mean supervectors of the speaker models are divided into several subvectors by mixture index. However, this division strategy cannot capture local acoustic class structure information among similar acoustic classes or discriminative information between different acoustic classes. In order to verify whether or not local structure information can help improve system performance, we develop five different speaker supervector segmentation methods. Experiments on NIST SRE08 prove that clustering similar acoustic classes together improves the system performance. In particular, the proposed method of equal size clustering achieves 5.1% relative decrease on EER compared to FSH1.
Na Li 0012, Weiwu Jiang, Helen M. Meng, Zhifeng Li 0001
ICASSP1