Han Zhu 0004

dblp:02/1555-4 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
9since 2021 · last 2025
0009-0002-5060-4454ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Hybrid Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
abstract
Pseudo-labeling based semi-supervised learning can mitigate the performance degradation resulting from the absence of labeled data in the target domain. In pseudo-labeling, the quality of pseudo-labels is crucial for the final performance. However, most works overlook the potential benefits of using decoder for pseudo-labels within the the mainstream hybrid Connectionist Temporal Classification (CTC) and attention (CTC/attention) based ASR architecture. Therefore, we propose Hybrid Pseudo-Labeling (HPL) to improve the quality of pseudo-labels during online decoding. HPL introduces a second-stage decoding using the decoder to alleviate substitution errors arising from the conditional independence assumption inherent in CTC for error correction. Furthermore, we propose Hybrid Selection to optimally combine results of encoder and decoder. Additionally, we introduce Speed Perturbation Enhancement (SPE) to further enhance the quality of pseudo-labels via speed perturbation. Experiments demonstrate that HPL achieves state-of-the-art performance compared to other mainstream pseudo-labeling methods.
Han Zhu 0004, Chengxu Yang, Gaofeng Cheng, Ta Li
ICASSP2
2025 CR-CTC: Consistency regularization on CTC for improved speech recognition
abstract
Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance. In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC. It significantly improves the CTC performance, achieving state-of-the-art results comparable to those attained by transducer or systems combining CTC and attention-based encoder-decoder (CTC/AED). We release our code at \url{https://github.com/k2-fsa/icefall}.
Zengwei Yao, Wei Kang 0006, Xiaoyu Yang 0005, Liyong Guo, Han Zhu 0004, Zengrui Jin, Zhaoqing Li, Long Lin, Daniel Povey
ICLR6
2024 Unsupervised Domain Adaptation on End-to-End Multi-Talker Overlapped Speech Recognition
abstract
Serialized Output Training (SOT) has emerged as the mainstream approach for addressing the multi-talker overlapped speech recognition challenge due to its simplicity. However, SOT encounters cross-domain performance degradation which hinders its application. Meanwhile, traditional domain adaption methods may harm the accuracy of speaker change point prediction evaluated by UD-CER, which is an important metric in SOT. To solve these issues, we propose Pseudo-Labeling based SOT (PL-SOT) for domain adaptation by treating speaker change token ($< $sc$>$) specially during training to increase the accuracy of speaker change point prediction. Firstly, we improve CTC loss by proposingWeakening and Enhancing CTC(WE-CTC) loss to weaken the learning of error-prone labels surrounding$<$sc$>$while enhance the emission probability of$< $sc$>$through modifying posteriors of the pseudo-labels. Secondly, we introduceWeighted Confidence Filter(WCF) that assigns higher scores of$<$sc$>$to exclude low-quality pseudo-labels without hurting the$< $sc$>$prediction. Experimental results show that PL-SOT achieves 17.7%/12.8% average relative reduction of CER/UD-CER, with AliMeeting as source domain and AISHELL-4 along with MagicData-RAMC as target domain.
Han Zhu 0004, Sanli Tian, Qingwei Zhao, Ta Li
IEEE Signal Process. Lett.2
2024 Boosting Cross-Domain Speech Recognition With Self-Supervision
abstract
The cross-domain performance of automatic speech recognition (ASR) could be severely hampered due to the mismatch between training and testing distributions. Since the target domain usually lacks labeled data, and domain shifts exist at acoustic and linguistic levels, it is challenging to perform unsupervised domain adaptation (UDA) for ASR. Previous work has shown that self-supervised learning (SSL) or pseudo-labeling (PL) is effective in UDA by exploiting the self-supervisions of unlabeled data. However, these self-supervisions also face performance degradation in mismatched domain distributions, which previous work fails to address. This work presents a systematic UDA framework to fully utilize the unlabeled data with self-supervision in the pre-training and fine-tuning paradigm. On the one hand, we apply continued pre-training and data replay techniques to mitigate the domain mismatch of the SSL pre-trained model. On the other hand, we propose a domain-adaptive fine-tuning approach based on the PL technique with three unique modifications: Firstly, we design a dual-branch PL method to decrease the sensitivity to the erroneous pseudo-labels; Secondly, we devise an uncertainty-aware confidence filtering strategy to improve pseudo-label correctness; Thirdly, we introduce a two-step PL approach to incorporate target domain linguistic knowledge, thus generating more accurate target domain pseudo-labels. Experimental results on various cross-domain scenarios demonstrate that the proposed approach effectively boosts the cross-domain performance and significantly outperforms previous approaches.
Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Wenxin Hou, Pengyuan Zhang, Yonghong Yan 0002
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
abstract
When labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, pseudo-labels are often noisy, containing numerous incorrect tokens. Taking noisy labels as ground-truth in the loss function results in suboptimal performance. Previous works attempted to mitigate this issue by either filtering out the nosiest pseudo-labels or improving the overall quality of pseudo-labels. While these methods are effective to some extent, it is unrealistic to entirely eliminate incorrect tokens in pseudo-labels. In this work, we propose a novel framework named alternative pseudo-labeling to tackle the issue of noisy pseudo-labels from the perspective of the training objective. The framework comprises several components. Firstly, a generalized CTC loss function is introduced to handle noisy pseudo-labels by accepting alternative tokens in the positions of incorrect tokens. Applying this loss function in pseudo-labeling requires detecting incorrect tokens in the predicted pseudo-labels. In this work, we adopt a confidence-based error detection method that identifies the incorrect tokens by comparing their confidence scores with a given threshold, thus necessitating the confidence score to be discriminative. Hence, the second proposed technique is the contrastive CTC loss function that widens the confidence gap between the correctly and incorrectly predicted tokens, thereby improving the error detection ability. Additionally, obtaining satisfactory performance with confidence-based error detection typically requires extensive threshold tuning. Instead, we propose an automatic thresholding method that uses labeled data as a proxy for determining the threshold, thus saving the pain of manual tuning. Experiments demonstrate that alternative pseudo-labeling outperforms existing pseudo-labeling approaches on datasets in various domains and languages.
Han Zhu 0004, Dongji Gao, Gaofeng Cheng, Daniel Povey, Pengyuan Zhang, Yonghong Yan 0002
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Wav2vec-S: Semi-Supervised Pre-Training for Low-Resource ASR
abstract
Self-supervised pre-training could effectively improve the performance of low-resource automatic speech recognition (ASR).However, existing self-supervised pre-training are taskagnostic, i.e., could be applied to various downstream tasks.Although it enlarges the scope of its application, the capacity of the pre-trained model is not fully utilized for the ASR task, and the learned representations may not be optimal for ASR.In this work, in order to build a better pre-trained model for low-resource ASR, we propose a pre-training approach called wav2vec-S, where we use task-specific semi-supervised pretraining to refine the self-supervised pre-trained model for the ASR task thus more effectively utilize the capacity of the pretrained model to generate task-specific representations for ASR.Experiments show that compared to wav2vec 2.0, wav2vec-S only requires a marginal increment of pre-training time but could significantly improve ASR performance on in-domain, cross-domain and cross-lingual datasets.Average relative WER reductions are 24.5% and 6.6% for 1h and 10h fine-tuning, respectively.Furthermore, we show that semi-supervised pretraining could close the representation gap between the selfsupervised pre-trained model and the corresponding fine-tuned model through canonical correlation analysis.
Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Pengyuan Zhang, Yonghong Yan 0002
INTERSPEECH1
2022 Decoupled Federated Learning for ASR with Non-IID Data
abstract
Automatic speech recognition (ASR) with federated learning (FL) makes it possible to leverage data from multiple clients without compromising privacy. The quality of FL-based ASR could be measured by recognition performance, communication and computation costs. When data among different clients are not independently and identically distributed (non-IID), the performance could degrade significantly. In this work, we tackle the non-IID issue in FL-based ASR with personalized FL, which learns personalized models for each client. Concretely, we propose two types of personalized FL approaches for ASR. Firstly, we adapt the personalization layer based FL for ASR, which keeps some layers locally to learn personalization models. Secondly, to reduce the communication and computation costs, we propose decoupled federated learning (DecoupleFL). On one hand, DecoupleFL moves the computation burden to the server, thus decreasing the computation on clients. On the other hand, DecoupleFL communicates secure high-level features instead of model parameters, thus reducing communication cost when models are large. Experiments demonstrate two proposed personalized FL-based ASR approaches could reduce WER by 2.3% - 3.4% compared with FedAvg. Among them, DecoupleFL has only 11.4% communication and 75% computation cost compared with FedAvg, which is also significantly less than the personalization layer based FL.
Han Zhu 0004, Jindong Wang 0001, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002
INTERSPEECH1
2022 Exploiting Adapters for Cross-Lingual Low-Resource Speech Recognition
abstract
Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource languages to build models for a low-resource target language. Since the low-resource language has limited training data, speech recognition models can easily overfit. Adapter is a versatile module that can be plugged into Transformer for parameter-efficient learning. In this paper, we propose to use adapters for parameter-efficient cross-lingual speech adaptation. Based on our previous MetaAdapter that implicitly leverages adapters, we propose a novel algorithm called SimAdapter for explicitly learning knowledge from adapters. Our algorithms can be easily integrated into the Transformer structure. MetaAdapter leverages meta-learning to transfer the general knowledge from training data to the test language. SimAdapter aims to learn the similarities between the source and target languages during fine-tuning using the adapters. We conduct extensive experiments on five-low-resource languages in the Common Voice dataset. Results demonstrate that MetaAdapter and SimAdapter can reduce WER by 2.98% and 2.55% with only 2.5% and 15.5% of trainable parameters compared to the strong full-model fine-tuning baseline. Moreover, we show that these two novel algorithms can be integrated for better performance with up to 3.55% relative WER reduction.
Wenxin Hou, Han Zhu 0004, Yidong Wang 0003, Jindong Wang 0001, Tao Qin 0001, Renjun Xu, Takahiro Shinozaki
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data
abstract
This paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encoder states. By this pre-training strategy, the decoder can learn how to generate grammatical text sequence before learning how to generate correct transcriptions. Contrast to other methods which utilize text only data to improve the ASR performance, our method does not change the network architecture of the ASR model or introduce extra component like text-to-speech (TTS) or text-to-encoder (TTE). Experimental results on LibriSpeech corpus show that the proposed method can relatively reduce the word error rate over 10%, using 960 hours transcriptions.
Changfeng Gao, Gaofeng Cheng, Runyan Yang, Han Zhu 0004, Pengyuan Zhang, Yonghong Yan 0002
ICASSP4
2020 Domain Adaptation Using Class Similarity for Robust Speech Recognition
abstract
When only limited target domain data is available, domain adaptation could be used to promote performance of deep neural network (DNN) acoustic model by leveraging well-trained source model and target domain data. However, suffering from domain mismatch and data sparsity, domain adaptation is very challenging. This paper proposes a novel adaptation method for DNN acoustic model using class similarity. Since the output distribution of DNN model contains the knowledge of similarity among classes, which is applicable to both source and target domain, it could be transferred from source to target model for the performance improvement. In our approach, we first compute the frame level posterior probabilities of source samples using source model. Then, for each class, probabilities of this class are used to compute a mean vector, which we refer to as mean soft labels. During adaptation, these mean soft labels are used in a regularization term to train the target model. Experiments showed that our approach outperforms fine-tuning using one-hot labels on both accent and noise adaptation task, especially when source and target domain are highly mismatched.
Han Zhu 0004, Jiangjiang Zhao, Yuling Ren, Pengyuan Zhang
INTERSPEECH1
2019 Multi-Accent Adaptation Based on Gate Mechanism
abstract
When only a limited amount of accented speech data is available, to promote multi-accent speech recognition performance, the conventional approach is accent-specific adaptation, which adapts the baseline model to multiple target accents independently. To simplify the adaptation procedure, we explore adapting the baseline model to multiple target accents simultaneously with multi-accent mixed data. Thus, we propose using accent-specific top layer with gate mechanism (AST-G) to realize multi-accent adaptation. Compared with the baseline model and accent-specific adaptation, AST-G achieves 9.8% and 1.9% average relative WER reduction respectively. However, in real-world applications, we can't obtain the accent category label for inference in advance. Therefore, we apply using an accent classifier to predict the accent label. To jointly train the acoustic model and the accent classifier, we propose the multi-task learning with gate mechanism (MTL-G). As the accent label prediction could be inaccurate, it performs worse than the accent-specific adaptation. Yet, in comparison with the baseline model, MTL-G achieves 5.1% average relative WER reduction.
Han Zhu 0004, Pengyuan Zhang, Yonghong Yan 0002
INTERSPEECH1