EDBT 2026 Demo / reviewers in the wild / expert
Yuke Lin
dblp:331/0832
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Codec-ASV: Exploring Neural Audio Codec For Speaker Representation LearningabstractDiscrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be largely preserved in the compression process since the reconstructed voice is almost the same in human listening. In this paper, we explore various training strategies and codec types for NAC-based speaker representation learning. Using ECAPA-TDNN as the model backbone, our approach achieves state-of-the-art performance with a 2.08% EER in NAC-based speaker verification scenarios. To better retain speaker information in early, more compressed layers, we introduce mask-layer augmentation and embedding fusion techniques during the training process. Experimental results show the effectiveness of our methods, particularly when inferring with limited codec layers. Yuke Lin, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026 |
ICASSP | 1 |
| 2025 | Robust Personal Voice Activity Detection for Mitigating Domain Mismatch and False Acceptance Scenarios
Yuke Lin, Longshuai Xiao, Chao Weng |
INTERSPEECH | 1 |
| 2024 | An Explainable and Lightweight Deep Learning Model with Attention Mechanism for Efficient Lung Disease DetectionabstractCOVID-19 has rapidly spread across the world as an extremely contagious disease. Therefore, early detection of the virus is essential to effectively control its transmission and prevent further spread. It is widely reported in the literature that convolutional neural networks (CNNs) are commonly utilized for COVID-19 detection. However, the majority of these models require substantial computational resources and possess a large number of parameters, rendering them less feasible for deployment on real-time devices. This study proposes a solution to address COVID-19 detection challenges by introducing LWIA-Net, a rapid and highly efficient lightweight CNN architecture. The LWIA-Net is equipped with channel attention squeeze and excitation (SE) blocks, as well as a naïve inception module, to enhance the network’s learning ability, reduce computational complexity, and maintain a high level of detection performance. By incorporating the naïve inception module, the LWIA-Net enables multi-level feature extraction, while the inclusion of the SE block helps focuses on informative channels. This integration contributes to the production of high-quality image features and effectively reduces redundancy. We performed thorough experiments using both our locally developed dataset and a publicly available dataset. The statistical analysis demonstrate that LWIA-Net achieves superior performance compared to pre-trained CNN architectures, despite its extremely lightweight architecture which demands low computational cost, memory space, and consists of only 1.11 million parameters. We conducted a comprehensive ablation study to gain a deeper understanding of the significance of each component and to verify their individual contributions to the LWIA-Net. Additionally, we utilized Grad-CAM analysis to confirm that the proposed model effectively identifies and emphasizes the most critical regions of interest. Experimental results have shown that LWIA-Net achieves remarkable accuracy rates of 94.45% and 97.56% on a local dataset and a public dataset. These findings suggest that LWIA-Net could potentially assist medical professionals in identifying patients with COVID-19 infection. Tingting Qian, Sohaib Asif, Yuke Lin, Jincao Yao, Enyu Wang, Vicky Yang Wang, Dong Xu 0006 |
BIBM | 3 |
| 2024 | Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker VerificationabstractUtilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-supervised domain adaptation. Firstly, we utilize limited labeled data from the target domain to derive domain-specific descriptors based on multiple distinct objectives, namely within-graph denoising, intra-class denoising and inter-class denoising. Then, the Infomap algorithm is adopted for embedding clustering, and the descriptors are leveraged to further refine the target domain’s pseudo-labels. Moreover, to further improve the quality of pseudo labels, we introduce the subcenter-purification and progressive-merging strategy for label denoising. Our proposed MoPC method achieves 4.95% EER and ranked the 1stplace on the evaluation set of VoxSRC 2023 track 3. We also conduct additional experiments on the FFSVC dataset and yield promising results. Ze Li 0003, Yuke Lin, Xiaoyi Qin, Haiying Wu, Ming Li 0026 |
ICASSP | 2 |
| 2024 | Voxblink: A Large Scale Speaker Verification Dataset on CameraabstractIn this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature of automated data collection, introducing noisy data is inevitable. Therefore, we also utilize a multi-modal purification step to generate a cleaner version of the VoxBlink, named VoxBlink-clean, comprising 18K identities and 1.02M utterances. In contrast to the VoxCeleb, the VoxBlink sources from short videos of ordinary users, and the covered scenarios can better align with real-life situations. To our best knowledge, the VoxBlink dataset is one of the largest publicly available speaker verification datasets. Leveraging the VoxCeleb and VoxBlink-clean datasets together, we employ diverse speaker verification models with multiple architectural backbones to conduct comprehensive evaluations on the VoxCeleb test sets. Experimental results indicate a substantial enhancement in performance—ranging from 12% to 30% relatively—across various backbone architectures upon incorporating the VoxBlink-clean into the training process. The details of the dataset can be found on $\color{Fuchsia} {{\text{Site}}}$. Yuke Lin, Xiaoyi Qin, Ming Cheng 0005, Haiying Wu, Ming Li 0026 |
ICASSP | 1 |
| 2024 | KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
Huali Zhou, Yuke Lin, Dong Liu 0028, Ming Li 0026 |
ICPR (23) | 2 |
| 2024 | VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Yuke Lin, Ming Cheng 0005, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li 0026 |
INTERSPEECH | 1 |
| 2024 | The Database and Benchmark For the Source Speaker Tracing Challenge 2024abstractVoice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/ Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026 |
SLT | 2 |
| 2023 | Haha-POD: An Attempt for Laughter-Based Non-Verbal Speaker VerificationabstractIt is widely acknowledged that discriminative representation for speaker verification can be extracted from verbal speech. However, how much speaker information that non-verbal vocalization carries is still a puzzle. This paper explores speaker verification based on the most ubiquitous form of non-verbal voice, laughter. First, we use a semi-automatic pipeline to collect a new Haha-Pod dataset from open-source podcast media. The dataset contains over 240 speakers’ laughter clips with corresponding high-quality verbal speech. Second, we propose a Two-Stage Teacher-Student (2S-TS) framework to minimize the within-speaker embedding distance between verbal and non-verbal (laughter) signals. Considering Haha-Pod as a test set, two trial sets (S2L-Eval) are designed to verify the speaker’s identity through laugh sounds. Experimental results demonstrate that our method can significantly improve the performance of the S2L-Eval test set with only a minor degradation on the VoxCeleb1 test set. The resources for the Haha-Pod dataset can be found at https://github.com/nevermoreLin/HahaPod. Yuke Lin, Xiaoyi Qin, Ming Li 0026 |
ASRU | 1 |