VLDB 2026 Research / reviewers in the wild / expert
Zhenduo Zhao
dblp:319/2350
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0002-7376-0441ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-Efficient Semi-Supervised Few-Shot Speaker Verification via Prototype Space OptimizationabstractSpeaker verification technology has widespread applications across many domains, benefiting from deep learning advancements. However, due to the high cost of acquiring labeled data, semi-supervised learning has emerged as a prominent research focus. Current semi-supervised learning frameworks commonly suffer from two limitations: (1) the labeled data distribution is often restricted, and (2) they still rely on a considerable amount of labeled data. To address these issues, we propose three different distribution scenarios of labeled data and construct a general semi-supervised framework. Furthermore, to enhance the guidance efficacy of limited labeled data, we innovatively employ prototype space optimization to strengthen the model's discriminative capability under low-resource scenarios. Experimental results demonstrate that on the Vox1-o test set, our approach achieves a 41.7% relative reduction in equal error rate compared to self-supervised baselines, and a 28.9% improvement over conventional semi-supervised framework baselines. Zhenduo Zhao, Shisong Wu, Pengyuan Zhang, Xueshuai Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Multi-Branch Coordinate Attention With Channel Dynamic Difference for Speaker VerificationabstractIn prior studies, researchers have proved the excellent performance of various deep neural networks on speaker verification (SV). However, most of the improvements of SV systems are aimed at modifying the specific network structure to enhance its robustness but with limited flexibility. In this paper, MCA-CDD, which is a novel universal residual block module and can easily replace the original residual block without increasing the block number, is proposed with adaptive multi-branch coordinate attention (MCA) and channel-level dynamic difference (CDD). The design of multiple branches enables CA to model time and frequency at multiple scales. In addition, CDD fusion is applied into the feature fusion process of the residual block. The CDD fusion at shallow positions of the model enables the model to learn the detailed speaker-related dynamic texture information of speech. Experiments are conducted on several ResNet backbones and the results on different seen and unseen test sets show significant improvements, outperforming the baseline by about relatively 15-20%. Zhenduo Zhao, Shisong Wu, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Progressive channel fusion for more efficient TDNN on speaker verification
Zhenduo Zhao |
Speech Commun. | 1 |
| 2024 | Prototype Division for Self-Supervised Speaker VerificationabstractSelf-supervised learning has shown promising performance on speaker verification tasks, among which Self DIstillation with NO labels (DINO) is currently a widely adopted framework. As one of the unsupervised deep clustering methods, the number of valid prototypes in DINO is far less than the speakers in practical applications and remains unchanged throughout the training period, leading to severe speaker confusion and performance degradation. Therefore, a strategy named prototype division (PD) is proposed to iteratively generate fine-grained prototypes in the projection space based on the converged model to separate confused categories, where new prototypes are derived from the neighborhood of the existing valid prototypes by clustering or sampling. The results on Vox1O achieve significant improvements, relatively outperforming the baseline by 31.1% without any auxiliary loss. Further experiments on CN-Celeb also show stable improvement, proving the consistency of the proposed method. Zhenduo Zhao, Zhuo Li 0020, Xueshuai Zhang, Pengyuan Zhang |
IEEE Signal Process. Lett. | 1 |
| 2023 | PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker VerificationabstractECAPA-TDNN is currently the most popular TDNN-series model for speaker verification, which refreshed the state-of-the-art (SOTA) performance of TDNN models. However, one-dimensional convolution has a global receptive field over the feature channel. It destroys the time-frequency relevance of the spectrogram. Besides, as ECAPA-TDNN only has five layers, a much shallower structure compared to ResNet restricts the capability to generate deep representations. To further improve ECAPA-TDNN, we propose a progressive channel fusion strategy that splits the spectrogram across the feature channel and gradually expands the receptive field through the network. Secondly, we enlarge the model by extending the depth and adding branches. Our proposed model achieves EER with 0.718 and minDCF(0.01) with 0.0858 on vox1o, relatively improved 16.1% and 19.5% compared with ECAPA-TDNN(C=1024). Zhenduo Zhao, Pengyuan Zhang |
ICASSP | 1 |
| 2023 | How to make embeddings suitable for PLDA
Zhuo Li 0020, Runqiu Xiao, Hangting Chen, Zhenduo Zhao, Pengyuan Zhang |
Comput. Speech Lang. | 4 |
| 2022 | The HCCL System for the NIST SRE21
Zhuo Li 0020, Runqiu Xiao, Hangting Chen, Zhenduo Zhao |
INTERSPEECH | 4 |