EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyi Qin
dblp:70/10621
· DBLP profile ↗
23ranked-venue papers
9as first author
15since 2021 · last 2024
0000-0003-2521-8084ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 8 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker VerificationabstractUtilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-supervised domain adaptation. Firstly, we utilize limited labeled data from the target domain to derive domain-specific descriptors based on multiple distinct objectives, namely within-graph denoising, intra-class denoising and inter-class denoising. Then, the Infomap algorithm is adopted for embedding clustering, and the descriptors are leveraged to further refine the target domain’s pseudo-labels. Moreover, to further improve the quality of pseudo labels, we introduce the subcenter-purification and progressive-merging strategy for label denoising. Our proposed MoPC method achieves 4.95% EER and ranked the 1stplace on the evaluation set of VoxSRC 2023 track 3. We also conduct additional experiments on the FFSVC dataset and yield promising results. Ze Li 0003, Yuke Lin, Xiaoyi Qin, Haiying Wu, Ming Li 0026 |
ICASSP | 4 |
| 2024 | Voxblink: A Large Scale Speaker Verification Dataset on CameraabstractIn this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature of automated data collection, introducing noisy data is inevitable. Therefore, we also utilize a multi-modal purification step to generate a cleaner version of the VoxBlink, named VoxBlink-clean, comprising 18K identities and 1.02M utterances. In contrast to the VoxCeleb, the VoxBlink sources from short videos of ordinary users, and the covered scenarios can better align with real-life situations. To our best knowledge, the VoxBlink dataset is one of the largest publicly available speaker verification datasets. Leveraging the VoxCeleb and VoxBlink-clean datasets together, we employ diverse speaker verification models with multiple architectural backbones to conduct comprehensive evaluations on the VoxCeleb test sets. Experimental results indicate a substantial enhancement in performance—ranging from 12% to 30% relatively—across various backbone architectures upon incorporating the VoxBlink-clean into the training process. The details of the dataset can be found on $\color{Fuchsia} {{\text{Site}}}$. Yuke Lin, Xiaoyi Qin, Ming Cheng 0005, Haiying Wu, Ming Li 0026 |
ICASSP | 2 |
| 2024 | Investigating Long-Term and Short-Term Time-Varying Speaker VerificationabstractThe performance of speaker verification systems can be adversely affected by time domain variations. However, limited research has been conducted on time-varying speaker verification due to the absence of appropriate datasets. This paper aims to investigate the impact of long-term and short-term time-varying in speaker verification and proposes solutions to mitigate these effects. For long-term speaker verification (i.e., cross-age speaker verification), we introduce an age-decoupling adversarial learning method to learn age-invariant speaker representation by mining age information from the VoxCeleb dataset. For short-term speaker verification, we collect the SMIIP-TimeVarying (SMIIP-TV) Dataset, which includes recordings at multiple time slots every day from 373 speakers for 90 consecutive days and other relevant meta information. Using this dataset, we analyze the time-varying of speaker embeddings and propose a novel but realistic time-varying speaker verification task, termed incremental sequence-pair speaker verification. This task involves continuous interaction between enrollment audios and a sequence of testing audios with the aim of improving performance over time. We introduce the template updating method to counter the negative effects over time, and then formulate the template updating processing as a Markov Decision Process and propose a template updating method based on deep reinforcement learning (DRL). The policy network of DRL is treated as an agent to determine if and how much should the template be updated. In summary, this paper releases our collected database, investigates both the long-term and short-term time-varying scenarios and provides insights and solutions into time-varying speaker verification. Xiaoyi Qin, Na Li 0012, Shufei Duan, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Haha-POD: An Attempt for Laughter-Based Non-Verbal Speaker VerificationabstractIt is widely acknowledged that discriminative representation for speaker verification can be extracted from verbal speech. However, how much speaker information that non-verbal vocalization carries is still a puzzle. This paper explores speaker verification based on the most ubiquitous form of non-verbal voice, laughter. First, we use a semi-automatic pipeline to collect a new Haha-Pod dataset from open-source podcast media. The dataset contains over 240 speakers’ laughter clips with corresponding high-quality verbal speech. Second, we propose a Two-Stage Teacher-Student (2S-TS) framework to minimize the within-speaker embedding distance between verbal and non-verbal (laughter) signals. Considering Haha-Pod as a test set, two trial sets (S2L-Eval) are designed to verify the speaker’s identity through laugh sounds. Experimental results demonstrate that our method can significantly improve the performance of the S2L-Eval test set with only a minor degradation on the VoxCeleb1 test set. The resources for the Haha-Pod dataset can be found at https://github.com/nevermoreLin/HahaPod. Yuke Lin, Xiaoyi Qin, Ming Li 0026 |
ASRU | 2 |
| 2023 | Target-Speaker Voice Activity Detection Via Sequence-to-Sequence PredictionabstractTarget-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large-scale speakers and predict high-resolution voice activities. Experimental results show that larger speaker capacity and higher output resolution can significantly reduce the diarization error rate (DER), which achieves the new state-of-the-art performance of 4.55% on the VoxConverse test set and 10.77% on Track 1 of the DIHARD-III evaluation set under the widely-used evaluation metrics. Ming Cheng 0005, Weiqing Wang 0004, Yucong Zhang, Xiaoyi Qin, Ming Li 0026 |
ICASSP | 4 |
| 2023 | Robust Multi-Channel Far-Field Speaker Verification Under Different In-Domain Data Availability ScenariosabstractThe popularity and application of smart home devices have made far-field speaker verification an urgent need. However, speaker verification performance is unsatisfactory under far-field environments despite its significant improvements enabled by deep neural networks (DNN). In this paper, we summarize our previous work and propose multiple training strategies and models for multi-channel far-field speaker verification with different in-domain data availability scenarios. The experiments are conducted on the FFSVC20 dataset, and we proposed the cross-device and cross-domain trials. We focus on single-channel and multi-channel speaker verification training based on the dataset. For single-channel speaker verification, considering the size of training data and availability of labels, we introduce three training scenarios and given our proposed training methods, including 1) given zero out-of-domain data and few in-domain labeled data; 2) given large-scale out-of-domain labeled data and few in-domain labeled data; 3) given large-scale out-of-domain labeled data and few in-domain unlabeled data. To this end, we propose a meta-learning approach, refined transfer learning methods, and semi-supervised learning for three scenarios, respectively. For multi-channel speaker verification, we first introduce two types of 3 dimension convolution (3D Conv) residual network (ResNet) models proposed in our previous works, including fully 3D ResNet and incorporating 3D Conv with 2D Conv ResNet (3D2D-ResNet). In this paper, we propose channel-wise 3D squeeze-and-excitation ResNet (C3DSE-ResNet) and spatial-wise 3D SE ResNet (S3DSE-ResNet) to further explore the channel dependencies and improve the 3D ConvNet performance. The results show that the proposed strategies and models can significantly boost performance under the far-field scenario. Xiaoyi Qin, Danwei Cai, Ming Li 0026 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Towards Lightweight Applications: Asymmetric Enroll-Verify Structure for Speaker VerificationabstractWith the development of deep learning, automatic speaker verification has made considerable progress over the past few years. However, to design a lightweight and robust system with limited computational resources is still a challenging problem. Traditionally, a speaker verification system is symmetrical, indicating that the same embedding extraction model is applied for both enrollment and verification in inference. In this paper, we come up with an innovative asymmetric structure, which takes the large-scale ECAPA-TDNN model for enrollment and the small-scale ECAPA-TDNNLite model for verification. As a symmetrical system, our proposed ECAPA-TDNNLite model achieves an EER of 3.07% on the Voxceleb1 original test set with only 11.6M FLOPS. Moreover, the asymmetric structure further reduces the EER to 2.31%, without increasing any computational costs during verification. Qingjian Li, Lin Yang 0014, Xiaoyi Qin, Junjie Wang 0010, Ming Li 0026 |
ICASSP | 4 |
| 2022 | Simple Attention Module Based Speaker Verification with Iterative Noisy Label DetectionabstractRecently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (SimAM), for speaker verification. The SimAM module is a plug-and-play module without extra modal parameters. In addition, we propose a noisy label detection method to iteratively filter out the data samples with a noisy label from the training data, considering that a large-scale dataset labeled with human annotation or other automated processes may contain noisy labels. Data with the noisy label may over parameterize a deep neural network (DNN) and result in a performance drop due to the memorization effect of the DNN. Experiments are conducted on VoxCeleb dataset. The speaker verification model with SimAM achieves the 0.675% equal error rate (EER) on VoxCeleb1 original test trials. Our proposed iterative noisy label detection method further reduces the EER to 0.643%. Xiaoyi Qin, Na Li 0012, Chao Weng, Dan Su 0002, Ming Li 0026 |
ICASSP | 1 |
| 2022 | Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for the M2met ChallengeabstractDukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based target-speaker voice activity detection (TS-VAD) to find the overlap between speakers. Firstly, we separately train a single-channel model for each of the 8 channels and fuse the results. In addition, we also employ the cross-channel self-attention to further improve the performance, where the non-linear spatial correlations between different channels are learned and fused. Experimental results on the evaluation set show that the single-channel TS-VAD reduces the DER by over 75% from 12.68% to 3.14%. The multi-channel TS-VAD further reduces the DER by 28% and achieves a DER of 2.26%. Our final submitted system achieves a DER of 2.98% on the AliMeeting test set, which ranks 1st in the M2MET challenge. In this challenge, our team is denoted as A41. Weiqing Wang 0004, Xiaoyi Qin, Ming Li 0026 |
ICASSP | 2 |
| 2022 | SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and MachinesabstractNowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people’s attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for speaker-content disentanglement of speech to better remove speaker information and get pure content information. Accordingly, our proposed framework contains a module that removes the speaker information from the acoustic feature of the source speaker. Moreover, speaker information control is added to our system to maintain the voice cloning performance. The proposed system is evaluated by subjective and objective metrics. Results show that our proposed system significantly reduces the trade-off problem in zero-shot voice conversion, while it also manages to have high spoofing power to the speaker verification system. Zexin Cai, Xiaoyi Qin, Ming Li 0026 |
ICASSP | 3 |
| 2022 | Cross-Age Speaker Verification: Learning Age-Invariant Speaker EmbeddingsabstractAutomatic speaker verification has achieved remarkable progress in recent years.However, there is little research on cross-age speaker verification (CASV) due to insufficient relevant data.In this paper, we mine cross-age test sets based on the VoxCeleb dataset and propose our age-invariant speaker representation(AISR) learning method.Since the VoxCeleb is collected from the YouTube platform, the dataset consists of crossage data inherently.However, the meta-data does not contain the speaker age label.Therefore, we adopt the face age estimation method to predict the speaker age value from the associated visual data, then label the audio recording with the estimated age.We construct multiple Cross-Age test sets on VoxCeleb (Vox-CA), which deliberately select the positive trials with large age-gap.Also, the effect of nationality and gender is considered in selecting negative pairs to align with Vox-H cases.The baseline system performance drops from 1.939% EER on the Vox-H test set to 10.419% on the Vox-CA20 test set, which indicates how difficult the cross-age scenario is.Consequently, we propose an age-decoupling adversarial learning (ADAL) method to alleviate the negative effect of the age gap and reduce intra-class variance.Our method outperforms the baseline system by over 10% related EER reduction on the Vox-CA20 test set.The source code and trial resources are available on https://github.com/qinxiaoyi/Cross-AgeSpeaker Verification. Xiaoyi Qin, Na Li 0012, Chao Weng, Dan Su 0002, Ming Li 0026 |
INTERSPEECH | 1 |
| 2022 | The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification ChallengeabstractThis paper describes our DKU-OPPO system for the 2022 Spoofing-Aware Speaker Verification (SASV) Challenge.First, we split the joint task into speaker verification (SV) and spoofing countermeasure (CM), these two tasks which are optimized separately.For ASV systems, four state-of-the-art methods are employed.For CM systems, we propose two methods on top of the challenge baseline to further improve the performance, namely Embedding Random Sampling Augmentation (ERSA) and One-Class Confusion Loss(OCCL).Second, we also explore whether SV embedding could help improve CM system performance.We observe a dramatic performance degradation of existing CM systems on the domain-mismatched Voxceleb2 dataset.Third, we compare different fusion strategies, including parallel score fusion and sequential cascaded systems.Compared to the 1.71% SASV-EER baseline, our submitted cascaded system obtains a 0.21% SASV-EER on the challenge official evaluation set. Xingming Wang, Xiaoyi Qin, Yikang Wang, Ming Li 0026 |
INTERSPEECH | 2 |
| 2021 | The 2020 Personalized Voice Trigger Challenge: Open Datasets, Evaluation Metrics, Baseline System and Results
Xingming Wang, Xiaoyi Qin, Yinping Zhang, Junjie Wang 0010, Dong Zhang 0002, Ming Li 0026 |
Interspeech | 3 |
| 2021 | Our Learned Lessons from Cross-Lingual Speaker Verification: The CRMI-DKU System Description for the Short-Duration Speaker Verification Challenge 2021
Xiaoyi Qin, Chao Wang 0111, Shilei Zhang, Ming Li 0026 |
Interspeech | 1 |
| 2021 | Binary Neural Network for Speaker VerificationabstractAlthough deep neural networks are successful for many tasks in the speech domain, the high computational and memory costs of deep neural networks make it difficult to directly deploy highperformance Neural Network systems on low-resource embedded devices. There are several mechanisms to reduce the size of the neural networks i.e. parameter pruning, parameter quantization, etc. This paper focuses on how to apply binary neural networks to the task of speaker verification. The proposed binarization of training parameters can largely maintain the performance while significantly reducing storage space requirements and computational costs. Experiment results show that, after binarizing the Convolutional Neural Network, the ResNet34-based network achieves an EER of around 5% on the Voxceleb1 testing dataset and even outperforms the traditional real number network on the text-dependent dataset: Xiaole while having a 32x memory saving. Tinglong Zhu, Xiaoyi Qin, Ming Li 0026 |
Interspeech | 2 |
| 2020 | HI-MIA: A Far-Field Text-Dependent Speaker Verification Database and the BaselinesabstractThis paper presents a far-field text-dependent speaker verification database named HI-MIA. We aim to meet the data requirement for far-field microphone array based speaker verification since most of the publicly available databases are single channel close-talking and text-independent. The database contains recordings of 340 people in rooms designed for the far-field scenario. Recordings are captured by multiple microphone arrays located in different directions and distance to the speaker and a high-fidelity close-talking microphone. Besides, we propose a set of end-to-end neural network based baseline systems that adopt single-channel data for training. Moreover, we propose a testing background aware enrollment augmentation strategy to further enhance the performance. Results show that the fusion systems could achieve 3.29% EER in the far-field enrollment far field testing task and 4.02% EER in the close-talking enrollment and far-field testing task. Xiaoyi Qin, Hui Bu, Ming Li 0026 |
ICASSP | 1 |
| 2020 | The INTERSPEECH 2020 Far-Field Speaker Verification ChallengeabstractThe INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field textindependent speaker verification from single microphone array, and far-field text-dependent speaker verification from distributed microphone arrays.All three tasks pose a cross-channel challenge to the participants.To simulate the real-life scenario, the enrollment utterances are recorded from close-talk cellphone, while the test utterances are recorded from the far-field microphone arrays.In this paper, we describe the database, the challenge, and the baseline system, which is based on a ResNetbased deep speaker network with cosine similarity scoring.For a given utterance, the speaker embeddings of different channels are equally averaged as the final embedding.The baseline system achieves minDCFs of 0.62, 0.66, and 0.64 and EERs of 6.27%, 6.55%, and 7.18% for task 1, task 2, and task 3, respectively. Xiaoyi Qin, Ming Li 0026, Hui Bu, Wei Rao 0002, Rohan Kumar Das, Shri Narayanan, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2019 | The DKU System for the Speaker Recognition Task of the 2019 VOiCES from a Distance ChallengeabstractIn this paper, we present the DKU system for the speaker recognition task of the VOiCES from a distance challenge 2019. We investigate the whole system pipeline for the far-field speaker verification, including data pre-processing, short-term spectral feature representation, utterance-level speaker modeling, back-end scoring, and score normalization. Our best single system employs a residual neural network trained with angular softmax loss. Also, the weighted prediction error algorithms can further improve performance. It achieves 0.3668 minDCF and 5.58% EER on the evaluation set by using a simple cosine similarity scoring. Finally, the submitted primary system obtains 0.3532 minDCF and 4.96% EER on the evaluation set. Danwei Cai, Xiaoyi Qin, Weicheng Cai, Ming Li 0026 |
INTERSPEECH | 2 |
| 2019 | Multi-Channel Training for End-to-End Speaker Recognition Under Reverberant and Noisy Environment
Danwei Cai, Xiaoyi Qin, Ming Li 0026 |
INTERSPEECH | 2 |
| 2019 | Polyphone Disambiguation for Mandarin Chinese Using Conditional Neural Network with Multi-Level Embedding FeaturesabstractThis paper describes a conditional neural network architecture for Mandarin Chinese polyphone disambiguation.The system is composed of a bidirectional recurrent neural network component acting as a sentence encoder to accumulate the context correlations, followed by a prediction network that maps the polyphonic character embeddings along with the conditions to corresponding pronunciations.We obtain the word-level condition from a pre-trained word-to-vector lookup table.One goal of polyphone disambiguation is to address the homograph problem existing in the front-end processing of Mandarin Chinese textto-speech system.Our system achieves an accuracy of 94.69% on a publicly available polyphonic character dataset.To further validate our choices on the conditional feature, we investigate polyphone disambiguation systems with multi-level conditions respectively.The experimental results show that both the sentence-level and the word-level conditional embedding features are able to attain good performance for Mandarin Chinese polyphone disambiguation. Zexin Cai, Yaogen Yang, Chuxiong Zhang, Xiaoyi Qin, Ming Li 0026 |
INTERSPEECH | 4 |
| 2019 | Far-Field End-to-End Text-Dependent Speaker Verification Based on Mixed Training Data with Transfer Learning and Enrollment Data Augmentation
Xiaoyi Qin, Danwei Cai, Ming Li 0026 |
INTERSPEECH | 1 |
| 2003 | An all-digital clock-smoothing technique - counting-prognosticationabstractThis article presents a novel universal all-digital clock-smoothing technique - counting-prognostication. Operation principles, performance analysis, and comparisons are given. Analysis and measurement results show that this technique can efficiently smooth jitter and wander for a wide pull-in range and pull-out range, and jitter accumulation is small. A cycle-varying counting-prognostication method, which decreases pull-in time, is also suggested. Xiaoyi Qin, Lieguang Zeng, Fuqin Xiong |
IEEE Trans. Commun. | 1 |
| 2003 | Coding, decoding, and recovery of clock synchronization in digital multiplexing systemabstractHigh-speed broadband digital communication networks rely on digital multiplexing technology where clock synchronization, including processing, transmission, and recovery of the clock, is the critical technique. This paper interprets the process of clock synchronization in multiplexing systems as quantizing and coding the information of clock synchronization, interprets clock justification as timing sigma-delta modulation (TΔ-ΣM), and interprets the jitter of justification as quantization error. As a result, decreasing the quantization error is equivalent to decreasing the jitter of justification. Using this theory, the paper studies the existing jitter-reducing techniques in transmitters and receivers, presents some techniques that can decrease the quantization error (justification jitter) in digital multiplexing systems, and presents a new method of clock recovery. Xiaoyi Qin, Lieguang Zeng, Fuqin Xiong |
IEEE Trans. Commun. | 2 |