VLDB 2026 Research / reviewers in the wild / expert
Danwei Cai
dblp:193/6521
· DBLP profile ↗
21ranked-venue papers
11as first author
10since 2021 · last 2024
0000-0002-5122-0623ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Joint Inference of Speaker Diarization and ASR with Multi-Stage Information SharingabstractIn this paper, we introduce a novel approach that unifies Automatic Speech Recognition (ASR) and speaker diarization in a cohesive framework. Utilizing the synergies between the two tasks, our method effectively extracts speaker-specific information from the lower layers of a pretrained Conformer-based ASR model while leveraging the higher layers for enhanced diarization performance. In particular, the integration of ASR contextual details into the diarization process has been demonstrated to be effective. Results on the DIHARD III dataset indicate that our approach achieves a Diarization Error Rate (DER) of 10.52%, which can be further reduced to 10.39% when integrating ASR features into the diarization model. These findings highlight the potential of our approach, suggesting competitive performance against other state-of-the-art systems. Additionally, our framework’s ability to simultaneously generate text transcripts for each speaker marks a distinct advantage, which can further enhance ASR capabilities and transition towards an end-to-end multitask framework encompassing both ASR and speaker diarization. Weiqing Wang 0004, Danwei Cai, Ming Cheng 0005, Ming Li 0026 |
ICASSP | 2 |
| 2024 | Leveraging ASR Pretrained Conformers for Speaker Verification Through Transfer Learning and Knowledge DistillationabstractThis paper focuses on the application of Conformers in speaker verification. Conformers, initially designed for Automatic Speech Recognition (ASR), excel at modeling both local and global contexts within speech signals effectively. Building on this synergistic relationship, this study introduces three strategies for leveraging ASR-pretrained Conformers in speaker verification: (1) Transfer learning: We use a pretrained ASR Conformer encoder to initialize the speaker embedding network, thereby enhancing model generalization and mitigating the risk of overfitting. (2) Knowledge distillation: We distill the complex capabilities of an ASR Conformer into a speaker verification model. This not only allows for flexibility in the student mode's network architecture but also incorporates frame-level ASR distillation loss as an auxiliary task to reinforce speaker verification. (3) Parameter-efficient transfer learning with speaker adaptation: A lightweight speaker adaptation module is proposed to convert ASR-derived features into speaker-specific embeddings, without altering the core architecture of the original ASR Conformer. This strategy facilitates the concurrent execution of ASR and speaker verification tasks within a singular model. Experiments were conducted on VoxCeleb datasets. The best model using the ASR pretraining method achieved a 0.43% equal error rate (EER) on the VoxCeleb1-O test trial, while the knowledge distillation approach yielded a 0.38% EER. Furthermore, by adding a mere 4.92 million parameters to a 130.94 million-parameter ASR Conformer encoder, the speaker adaptation approach achieved a 0.45% EER, enabling parallel speech recognition and speaker verification within a single ASR Conformer encoder. Overall, our techniques successfully transfer rich ASR knowledge to advanced speaker modeling. Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification SystemsabstractAn automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person’s speech signal to make it sound like another speaker’s voice and deceive the speaker verification system. Most countermeasures for voice conversion-based spoofing attacks are designed to discriminate bona fide speech from spoofed speech for speaker verification systems. In this paper, we investigate the problem of source speaker identification – inferring the identity of the source speaker given the voice converted speech. To perform source speaker identification, we simply add voice-converted speech data with the label of source speaker identity to the genuine speech dataset during speaker embedding network training. Experimental results show the feasibility of source speaker identification when training and testing with converted speeches from the same voice conversion model(s). In addition, our results demonstrate that having more converted utterances from various voice conversion model for training helps improve the source speaker identification performance on converted utterances from unseen voice conversion models. Danwei Cai, Zexin Cai, Ming Li 0026 |
ICASSP | 1 |
| 2023 | Pretraining Conformer with ASR for Speaker VerificationabstractThis paper proposes to pretrain Conformer with automatic speech recognition (ASR) task for speaker verification. Conformer combines convolution neural network (CNN) and Transformer model for modeling local and global features, respectively. Recently, multi-scale feature aggregation Conformer (MFA-Conformer) has been proposed for automatic speaker verification. MFA-Conformer concatenates frame-level outputs from all Conformer blocks for further pooling. However, our experiments show that Conformer can be easily overfitted with limited speaker recognition training data. To avoid overfitting, we propose to transfer the knowledge learned from ASR to speaker verification. Specifically, an ASR pretrained Conformer is used to initialize the training of MFA-Conformer for speaker verification. Our experiments show that pretraining Conformer with ASR leads to significant performance gains across model sizes. The best model achieves 0.48%, 0.71% and 1.54% EER on Voxceleb1-O, Voxceleb1-E, and Voxceleb1-H, respectively. Danwei Cai, Weiqing Wang 0004, Ming Li 0026, Chuanzeng Huang |
ICASSP | 1 |
| 2023 | Robust Multi-Channel Far-Field Speaker Verification Under Different In-Domain Data Availability ScenariosabstractThe popularity and application of smart home devices have made far-field speaker verification an urgent need. However, speaker verification performance is unsatisfactory under far-field environments despite its significant improvements enabled by deep neural networks (DNN). In this paper, we summarize our previous work and propose multiple training strategies and models for multi-channel far-field speaker verification with different in-domain data availability scenarios. The experiments are conducted on the FFSVC20 dataset, and we proposed the cross-device and cross-domain trials. We focus on single-channel and multi-channel speaker verification training based on the dataset. For single-channel speaker verification, considering the size of training data and availability of labels, we introduce three training scenarios and given our proposed training methods, including 1) given zero out-of-domain data and few in-domain labeled data; 2) given large-scale out-of-domain labeled data and few in-domain labeled data; 3) given large-scale out-of-domain labeled data and few in-domain unlabeled data. To this end, we propose a meta-learning approach, refined transfer learning methods, and semi-supervised learning for three scenarios, respectively. For multi-channel speaker verification, we first introduce two types of 3 dimension convolution (3D Conv) residual network (ResNet) models proposed in our previous works, including fully 3D ResNet and incorporating 3D Conv with 2D Conv ResNet (3D2D-ResNet). In this paper, we propose channel-wise 3D squeeze-and-excitation ResNet (C3DSE-ResNet) and spatial-wise 3D SE ResNet (S3DSE-ResNet) to further explore the channel dependencies and improve the 3D ConvNet performance. The results show that the proposed strategies and models can significantly boost performance under the far-field scenario. Xiaoyi Qin, Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Incorporating Visual Information in Audio Based Self-Supervised Speaker RecognitionabstractThe currentsuccess of deep learning largely benefits from the availability of large amount of labeled data. However, collecting a large-scale dataset with human annotation can be expensive and sometimes difficult. Self-supervised learning thus attracts many research interests to train models without labels. In this paper, we propose a self-supervised learning framework for speaker recognition. Combining clustering with deep representation learning, the proposed framework generates pseudo labels for the unlabeled dataset and learns speaker representation without human annotation. Our method starts with training a speaker representation encoder with contrastive self-supervised learning. Clustering on the learned representation generates pseudo labels, which are used as the supervisory signal for the subsequent training of the representation encoder. The clustering and representation learning process is performed iteratively to bootstrap the discriminative power of the deep neural network. We apply this self-supervised learning framework to both single modal audio data and multi-modal audio-visual data. For audio-visual data, audio and visual representation encoders are employed to learn representations of the corresponding modality. A cluster ensemble algorithm is then used to fuse the clustering results of the two modalities. The complementary information in multi-modalities ensures a robust and fault-tolerant supervisory signal for audio and visual representation learning. Experimental results show that our proposed iterative self-supervised learning framework outperforms previous works with self-supervision by large margins. Training with single modal audio data on the development set of VoxCeleb 2, our proposed framework achieves an equal error rate (EER) of 2.8% on the original test trials of VoxCeleb 1. When training with additional visual modality, the EER further reduces to 1.8%, which is only 20% higher than the fully supervised audio-based system with an EER of 1.5%. Also, experimental analysis shows that the proposed framework generates pseudolabels that are highly correlated to ground truth labels. Danwei Cai, Weiqing Wang 0004, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Similarity Measurement of Segment-Level Speaker Embeddings in Speaker DiarizationabstractIn this paper, we propose a neural-network-based similarity measurement method to learn the similarity between any two speaker embeddings, where both previous and future contexts are considered. Moreover, we propose the segmental pooling strategy and jointly train the speaker embedding network along with the similarity measurement model. Later, this joint training framework is further extended to the target-speaker voice activity detection (TS-VAD), with only slight modification in the network architecture. Experimental results of the DIHARD II, DIHARD III and VoxConverse datasets show that our clustering-based system with the neural similarity measurement achieves superior performance to recent approaches on all three datasets. In addition, the segment-level TS-VAD method further improves the clustering-based results and achieves DER of 16.48%, 11.62% and 4.39% on the DIHARD II, DIHARD III and VoxConverse datasets, respectively. Weiqing Wang 0004, Qingjian Lin, Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | An Iterative Framework for Self-Supervised Deep Speaker Representation LearningabstractIn this paper, we propose an iterative framework for self-supervised speaker representation learning based on a deep neural network (DNN). The framework starts with training a self-supervision speaker embedding network by maximizing agreement between different segments within an utterance via a contrastive loss. Taking advantage of DNN’s ability to learn from data with label noise, we propose to cluster the speaker embedding obtained from the previous speaker network and use the subsequent class assignments as pseudo labels to train a new DNN. Moreover, we iteratively train the speaker network with pseudo labels generated from the previous step to bootstrap the discriminative power of a DNN. Speaker verification experiments are conducted on the VoxCeleb dataset. The results show that our proposed iterative self-supervised learning framework outperformed previous works using self-supervision. The speaker network after 5 iterations obtains a 61% performance gain over the speaker embedding model trained with contrastive loss. Danwei Cai, Weiqing Wang 0004, Ming Li 0026 |
ICASSP | 1 |
| 2021 | The DKU-Duke-Lenovo System Description for the Fearless Steps Challenge Phase III
Weiqing Wang 0004, Danwei Cai, Qingjian Lin, Mi Hong, Ming Li 0026 |
Interspeech | 2 |
| 2021 | Embedding Aggregation for Far-Field Speaker Verification with Distributed Microphone ArraysabstractWith the successful application of deep speaker embedding networks, the performance of speaker verification systems has significantly improved under clean and close-talking settings; however, unsatisfactory performance persists under noisy and far-field environments. This study aims at improving the performance of far-field speaker verification systems with distributed microphone arrays in the smart home scenario. The proposed learning framework consists of two modules: a deep speaker embedding module and an aggregation module. The former extracts a speaker embedding for each recording. The latter, based on either averaged pooling or attentive pooling, aggregates speaker embeddings and learns a unified representation for all recordings captured by distributed microphone arrays. The two modules are trained in an end-to-end manner. To evaluate this framework, we conduct experiments on the real text-dependent far-field datasets Hi Mia. Results show that our framework outperforms the naive averaged aggregation methods by 20% in terms of equal error rate (EER) with six distributed microphone arrays. Also, we find that the attention-based aggregation advocates high-quality recordings and repels low-quality ones. Danwei Cai, Ming Li 0026 |
SLT | 1 |
| 2020 | Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy EnvironmentsabstractDespite the significant improvements in speaker recognition enabled by deep neural networks, unsatisfactory performance persists under noisy environments. In this paper, we train the speaker embedding network to learn the "clean" embedding of the noisy utterance. Specifically, the network is trained with the original speaker identification loss with an auxiliary within-sample variability-invariant loss. This auxiliary variability-invariant loss is used to learn the same embedding among the clean utterance and its noisy copies and prevents the network from encoding the undesired noises or variabilities into the speaker representation. Furthermore, we investigate the data preparation strategy for generating clean and noisy utterance pairs on-the-fly. The strategy generates different noisy copies for the same clean utterance at each training step, helping the speaker embedding network generalize better under noisy environments. Experiments on VoxCeleb1 indicate that the proposed training framework improves the performance of the speaker verification system in both clean and noisy conditions. Danwei Cai, Weicheng Cai, Ming Li 0026 |
ICASSP | 1 |
| 2019 | Utterance-level End-to-end Language Identification Using Attention-based CNN-BLSTMabstractIn this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level, which means the utterance-level decision can be directly obtained from the output of the neural network. To handle speech utterances with entire arbitrary and potentially long duration, we combine CNN-BLSTM model with a self-attentive pooling layer together. The front-end CNN-BLSTM module plays a role as local pattern extractor for the variable-length inputs, and the following self-attentive pooling layer is built on top to get the fixed-dimensional utterance-level representation. We conducted experiments on NIST LRE07 closed-set task, and the results reveal that the proposed attention-based CNN-BLSTM model achieves comparable error reduction with other state-of-the-art utterance-level neural network approaches for all 3 seconds, 10 seconds, 30 seconds duration tasks. Weicheng Cai, Danwei Cai, Shen Huang, Ming Li 0026 |
ICASSP | 2 |
| 2019 | The DKU-SMIIP System for NIST 2018 Speaker Recognition EvaluationabstractIn this paper, we present the system submission for the NIST 2018 Speaker Recognition Evaluation by DKU Speech and Multi-Modal Intelligent Information Processing (SMIIP) Lab.We explore various kinds of state-of-the-art front-end extractors as well as back-end modeling for text-independent speaker verifications.Our submitted primary systems employ multiple state-of-the-art front-end extractors, including the MFCC i-vector, the DNN tandem i-vector, the TDNN x-vector, and the deep ResNet.After speaker embedding is extracted, we exploit several kinds of back-end modeling to perform variability compensation and domain adaptation for mismatch training and testing conditions.The final submitted system on the fixed condition obtains actual detection cost of 0.392 and 0.494 on CMN2 and VAST evaluation data respectively.After the official evaluation, we further extend our experiments by investigating multiple encoding layer designs and loss functions for the deep ResNet system. Danwei Cai, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 1 |
| 2019 | The DKU System for the Speaker Recognition Task of the 2019 VOiCES from a Distance ChallengeabstractIn this paper, we present the DKU system for the speaker recognition task of the VOiCES from a distance challenge 2019. We investigate the whole system pipeline for the far-field speaker verification, including data pre-processing, short-term spectral feature representation, utterance-level speaker modeling, back-end scoring, and score normalization. Our best single system employs a residual neural network trained with angular softmax loss. Also, the weighted prediction error algorithms can further improve performance. It achieves 0.3668 minDCF and 5.58% EER on the evaluation set by using a simple cosine similarity scoring. Finally, the submitted primary system obtains 0.3532 minDCF and 4.96% EER on the evaluation set. Danwei Cai, Xiaoyi Qin, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 1 |
| 2019 | Multi-Channel Training for End-to-End Speaker Recognition Under Reverberant and Noisy Environment
Danwei Cai, Xiaoyi Qin, Ming Li 0026 |
INTERSPEECH | 1 |
| 2019 | The DKU Replay Detection System for the ASVspoof 2019 Challenge: On Data Augmentation, Feature Representation, Classification, and FusionabstractThis paper describes our DKU replay detection system for the ASVspoof 2019 challenge.The goal is to develop spoofing countermeasure for automatic speaker recognition in physical access scenario.We leverage the countermeasure system pipeline from four aspects, including the data augmentation, feature representation, classification, and fusion.First, we introduce an utterance-level deep learning framework for antispoofing.It receives the variable-length feature sequence and outputs the utterance-level scores directly.Based on the framework, we try out various kinds of input feature representations extracted from either the magnitude spectrum or phase spectrum.Besides, we also perform the data augmentation strategy by applying the speed perturbation on the raw waveform.Our best single system employs a residual neural network trained by the speed-perturbed group delay gram.It achieves EER of 1.04% on the development set, as well as EER of 1.08% on the evaluation set.Finally, using the simple average score from several single systems can further improve the performance.EER of 0.24% on the development set and 0.66% on the evaluation set is obtained for our primary system. Weicheng Cai, Haiwei Wu, Danwei Cai, Ming Li 0026 |
INTERSPEECH | 3 |
| 2019 | Survey Talk: End-to-End Deep Neural Network Based Speaker and Language Recognition
Ming Li 0026, Weicheng Cai, Danwei Cai |
INTERSPEECH | 3 |
| 2019 | Far-Field End-to-End Text-Dependent Speaker Verification Based on Mixed Training Data with Transfer Learning and Enrollment Data Augmentation
Xiaoyi Qin, Danwei Cai, Ming Li 0026 |
INTERSPEECH | 2 |
| 2018 | Cancellable speech template via random binary orthogonal matrices projection hashing
Kong-Yik Chee, Zhe Jin 0001, Danwei Cai, Ming Li 0026, Wun-She Yap, Yen-Lung Lai, Bok-Min Goi |
Pattern Recognit. | 3 |
| 2017 | Countermeasures for Automatic Speaker Verification Replay Spoofing Attack : On Data Augmentation, Feature Representation, Classification and Fusion
Weicheng Cai, Danwei Cai, Wenbo Liu 0002, Ming Li 0026 |
INTERSPEECH | 2 |
| 2017 | End-to-End Deep Learning Framework for Speech Paralinguistics Detection Based on Perception Aware Spectrum
Danwei Cai, Zhidong Ni, Wenbo Liu 0002, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 1 |