Youzhi Tu

dblp:243/6873 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0002-9580-2414ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Class unbiasing for generalization in medical diagnosis
Lishi Zuo, Lu Yi 0001, Youzhi Tu, Man-Wai Mak
Pattern Recognit.3
2025 Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition
abstract
Recent works suggest that decoupling the information of non-target speakers from that of the target speaker in knowledge distillation (KD) and subsequently emphasizing the former can lead to significant performance improvement. However, a well-trained teacher model typically produces almost zero non-target speaker posteriors with limited contribution to knowledge transfer, resulting in a less effective KD. To address this problem, we advocate a dual-group knowledge distillation framework, wherein the primary group with top-k speaker posteriors captures most of the speaker discrimination knowledge in an utterance. The non-primary group contributes to the KD through a binary classification (distillation) between the primary and non-primary groups. In addition, adaptive logit softening is proposed to adjust the teacher’s and student’s logits in the binary distillation, further facilitating effective knowledge transfer. The proposed method trained with a simple x-vector pipeline obtains an impressive equal error rate of 1.46%, 1.47%, and 2.70% on three VoxCeleb1 test sets, outperforming the state-of-the-art methods with a noticeable margin.
Chong-Xin Gan, Youzhi Tu, Zezhong Jin, Man-Wai Mak, Kong-Aik Lee
ICASSP2
2025 Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification
abstract
In recent years, there has been a surge in the use of a pre-trained speech model as a feature extractor for speaker verification (SV). To reduce model complexity, researchers transfer knowledge from a pre-trained model to a lightweight student model, enabling the latter to reach a performance level not attainable by conventional methods. However, due to the differences in model capacity, the student features contain more noise. This results in discrepancies between the teacher and student features at the intermediate layers, negatively impacting feature-level knowledge distillation (KD). To address this issue, we employ a diffusion model to denoise the student features for KD (DenoKD). This approach enables more effective feature-level distillation. Our method, trained with a small ECAPA-TDNN, achieved a 13% improvement over the baseline on the VoxCeleb1-O test set. Further more, the DenoKD mechanism is found to be effective for SV on short test utterances.
Zezhong Jin, Youzhi Tu, Zhe Li 0030, Chong-Xin Gan, Man-Wai Mak
ICASSP2
2025 Adversarially adaptive temperatures for decoupled knowledge distillation with applications to speaker verification
abstract
202502 bcch
Zezhong Jin, Youzhi Tu, Chong-Xin Gan, Man-Wai Mak, Kong-Aik Lee
Neurocomputing2
2025 ConFusionformer: Locality-enhanced Conformer through multi-resolution attention fusion for speaker verification
Youzhi Tu, Man-Wai Mak, Kong-Aik Lee, Weiwei Lin 0002
Neurocomputing1
2024 Contrastive Speaker Embedding With Sequential Disentanglement
abstract
Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional SimCLR framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that only the speaker factors are used for constructing a contrastive loss objective. Because content factors have been removed from the contrastive learning, the resulting speaker embeddings will be content-invariant. Experimental results on VoxCeleb1-test show that the proposed method consistently outperforms SimCLR. This suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
ICASSP1
2024 Promoting Independence of Depression and Speaker Features for Speaker Disentanglement in Speech-Based Depression Detection
abstract
Recent studies have demonstrated the effectiveness of speaker disentanglement in mitigating the interference caused by speaker features in speech-based depression detection. However, the inherent entanglement between depression features and speaker features poses challenges to depression detection. In this study, we propose a mutual information-based speaker-invariant depression detector (MI-SIDD) that aims to promote independence between depression and speaker features to facilitate speaker disentanglement. Specifically, we disentangle the speaker features using a vanilla autoencoder with a well-tuned bottleneck layer and minimize the mutual information between depression and speaker features using a conditional mutual information constraint. Experimental results demonstrate the effectiveness of speaker disentanglement and the promotion of independence between depression and speaker features. Our MI-SIDD model achieves competitive performance compared to state-of-the-art methods on the DAIC-WOZ dataset.
Lishi Zuo, Man-Wai Mak, Youzhi Tu
ICASSP3
2024 W-GVKT: Within-Global-View Knowledge Transfer for Speaker Verification
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Zezhong Jin, Youzhi Tu, Man-Wai Mak
INTERSPEECH2
2024 Self-Supervised Learning with Multi-Head Multi-Mode Knowledge Distillation for Speaker Verification
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Zezhong Jin, Youzhi Tu, Man-Wai Mak
INTERSPEECH2
2024 Contrastive Self-Supervised Speaker Embedding With Sequential Disentanglement
abstract
Contrastive self-supervised learning has been widely used in speaker embedding to address the labeling challenge. Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional contrastive learning framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that the speaker factors become the main contributor to the contrastive loss. Because content factors have been removed from contrastive learning, the resulting speaker embeddings will be content-invariant. The learned embeddings are also robust to language mismatch. It is shown that the proposed method consistently outperforms the conventional contrastive speaker embedding on the VoxCeleb1 and CN-Celeb datasets. This finding suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations
abstract
Self-supervised learning (SSL) speech models such as wav2vec and HuBERT have demonstrated state-of-the-art performance on automatic speech recognition (ASR) and proved to be extremely useful in low label-resource settings. However, the success of SSL models has yet to transfer to utterance-level tasks such as speaker, emotion, and language recognition, which still require supervised fine-tuning of the SSL models to obtain good performance. We argue that the problem is caused by the lack of disentangled representations and an utterance-level learning objective for these tasks. Inspired by how HuBERT uses clustering to discover hidden acoustic units, we formulate a factor analysis (FA) model that uses the discovered hidden acoustic units to align the SSL features. The underlying utterance-level representations are disentangled using probabilistic inference on the aligned features. Furthermore, the variational lower bound derived from the FA model provides an utterance-level objective, allowing error gradients to be backpropagated to the Transformer layers to learn highly discriminative acoustic units. When used in conjunction with HuBERT's masked prediction training, our models outperform the current best model, WavLM, on all utterance-level non-semantic tasks on the SUPERB benchmark with only 20% of labeled data.
Weiwei Lin 0002, Chenhang He, Man-Wai Mak, Youzhi Tu
ICML4
2022 Aggregating Frame-Level Information in the Spectral Domain With Self-Attention for Speaker Embedding
abstract
Most pooling methods in state-of-the-art speaker embedding networks are implemented in the temporal domain. However, due to the high non-stationarity in the feature maps produced from the last frame-level layer, it is not advantageous to use the global statistics (e.g., means and standard deviations) of the temporal feature maps as aggregated embeddings. This motivates us to explore stationary spectral representations and perform aggregation in the spectral domain. In this paper, we propose attentive short-time spectral pooling (attentive STSP) from a Fourier perspective to exploit the local stationarity of the feature maps. In attentive STSP, for each utterance, we compute the spectral representations through a weighted average of the windowed segments within each spectrogram by attention weights and aggregate their lowest spectral components to form the speaker embedding. Because most of the feature map energy is concentrated in the low-frequency region of the spectral domain, attentive STSP facilitates the information aggregation by retaining the low spectral components only. Attentive STSP is shown to consistently outperform attentive pooling on VoxCeleb1, VOiCES19-eval, SRE16-eval, and SRE18-CMN2-eval. This observation suggests that applying segment-level attention and leveraging low spectral components can produce discriminative speaker embeddings.
Youzhi Tu, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Short-Time Spectral Aggregation for Speaker Embedding
abstract
State-of-the-art speaker verification systems take frame-level acoustics features as input and produce fixed-dimensional embeddings as utterance-level representations. Thus, how to aggregate information from frame-level features is vital for achieving high performance. This paper introduces short-time spectral pooling (STSP) for better aggregation of frame-level information. STSP transforms the temporal feature maps of a speaker embedding network into the spectral domain and extracts the lowest spectral components of the averaged spectrograms for aggregation. Benefiting from the low-pass characteristic of the averaged spectrograms, STSP is able to preserve most of the speaker information in the feature maps using a few spectral components only. We show that statistics pooling is a special case of STSP where only the DC spectral components are used. Experiments on VoxCeleb1 and VOiCES 2019 show that STSP outperforms statistics pooling and multi-head attentive pooling, which suggests that leveraging more spectral information in the CNN feature maps can produce highly discriminative speaker embeddings.
Youzhi Tu, Man-Wai Mak
ICASSP1
2021 Mutual Information Enhanced Training for Speaker Embedding
abstract
22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021, Brno, Czechia, August 30 - September 3, 2021
Youzhi Tu, Man-Wai Mak
Interspeech1
2020 Information Maximized Variational Domain Adversarial Learning for Speaker Verification
abstract
Domain mismatch is a common problem in speaker verification. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) to reduce domain mismatch by incorporating an InfoVAE into domain adversarial training (DAT). DAT aims to produce speaker discriminative and domain-invariant features. The InfoVAE has two roles. First, it performs variational regularization on the learned features so that they follow a Gaussian distribution, which is essential for the standard PLDA backend. Second, it preserves mutual information between the features and the training set to extract extra speaker discriminative information. Experiments on both SRE16 and SRE18-CMN2 show that the InfoVDANN outperforms the recent VDANN, which suggests that increasing the mutual information between the latent features and input features enables the InfoVDANN to extract extra speaker information that is otherwise not possible.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
ICASSP1
2020 Variational Domain Adversarial Learning With Mutual Information Maximization for Speaker Verification
abstract
Domain mismatch is a common problem in speaker verification (SV) and often causes performance degradation. For the system relying on the Gaussian PLDA backend to suppress the channel variability, the performance would be further limited if there is no Gaussianity constraint on the learned embeddings. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) that incorporates an InfoVAE into domain adversarial training (DAT) to reduce domain mismatch and simultaneously meet the Gaussianity requirement of the PLDA backend. Specifically, DAT is applied to produce speaker discriminative and domain-invariant features, while the InfoVAE performs variational regularization on the embedded features so that they follow a Gaussian distribution. Another benefit of the InfoVAE is that it avoids posterior collapse in VAEs by preserving the mutual information between the embedded features and the training set so that extra speaker information can be retained in the features. Experiments on both SRE16 and SRE18-CMN2 show that the InfoVDANN outperforms the recent VDANN, which suggests that increasing the mutual information between the embedded features and input features enables the InfoVDANN to extract extra speaker information that is otherwise not possible.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Semi-supervised Nuisance-attribute Networks for Domain Adaptation
abstract
How to overcome the training and test data mismatch in speaker verification systems has been a focus of research recently. In this paper, we propose a semi-supervised nuisance attribute network (SNAN) to reduce the domain mismatch in i-vectors and x-vectors. SNANs are based on the idea of nuisance attribute removal in inter-dataset variability compensation (IDVC). But instead of measuring the domain variability through the dataset means, SNANs use the maximum mean discrepancy (MMD) as part of their loss function, which enables the network to find nuisance directions in which domain variability is measured up to infinite moment. The architecture of SNANs also allows us to incorporate the out-of-domain speaker labels into the semi-supervised training process through the center loss and triplet loss. Using SNANs as a preprocessing step for PLDA training, we achieve a relative improvement of 11.8% in EER on NIST 2016 SRE compared to PLDA without adaptation. We also found that the semi-supervised approach can further improve SNANs' performance.
Weiwei Lin 0002, Man-Wai Mak, Youzhi Tu, Jen-Tzung Chien
ICASSP3
2019 Variational Domain Adversarial Learning for Speaker Verification
abstract
20th Annual Conference of the International Speech Communication Association: Crossroads of Speech and Language, INTERSPEECH 2019, Graz, Austria, 15-19 September 2019
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
INTERSPEECH1