Wei Zhou 0043

dblp:69/5011-43 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
14since 2021 · last 2024
0009-0006-3754-8872ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021
YearPublicationVenuePosition
2024 On the Relation Between Internal Language Model and Sequence Discriminative Training for Neural Transducers
abstract
Internal language model (ILM) subtraction has been widely applied to improve the performance of the RNN-Transducer with external language model (LM) fusion for speech recognition. In this work, we show that sequence discriminative training has a strong correlation with ILM subtraction from both theoretical and empirical points of view. Theoretically, we derive that the global optimum of maximum mutual information (MMI) training shares a similar formula as ILM subtraction. Empirically, we show that ILM subtraction and sequence discriminative training achieve similar effects across a wide range of experiments on Librispeech, including both MMI and minimum Bayes risk (MBR) criteria, as well as neural transducers and LMs of both full and limited context. The benefit of ILM subtraction also becomes much smaller after sequence discriminative training. We also provide an indepth study to show that sequence discriminative training has a minimal effect on the commonly used zero-encoder ILM estimation, but a joint effect on both encoder and prediction + joint network for posterior probability reshaping including both ILM and blank suppression.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ICASSP2
2024 Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
Jingjing Xu 0002, Wei Zhou 0043, Eugen Beck, Ralf Schlüter
INTERSPEECH2
2023 Investigating The Effect of Language Models in Sequence Discriminative Training For Neural Transducers
abstract
In this work, we investigate the effect of language models (LMs) with different context lengths and label units (phoneme vs. word) used in sequence discriminative training for phoneme-based neural transducers. Both lattice-free and N-best-list approaches are examined. For lattice-free methods with phoneme-level LMs, we propose a method to approximate the context history to employ LMs with full-context dependency. This approximation can be extended to arbitrary context length and enables the usage of word-level LMs in lattice-free methods. Moreover, a systematic comparison is conducted across lattice-free and N-best-list-based methods. Experimental results on Librispeech show that using the word-level LM in training outperforms the phoneme-level LM. Besides, we find that the context size of the LM used for probability computation has a limited effect on performance. Moreover, our results reveal the pivotal importance of the hypothesis space quality in sequence discriminative training.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ASRU2
2023 Lattice-Free Sequence Discriminative Training for Phoneme-Based Neural Transducers
abstract
Recently, RNN-Transducers have achieved remarkable results on various automatic speech recognition tasks. However, lattice-free sequence discriminative training methods, which obtain superior performance in hybrid models, are rarely investigated in RNN-Transducers. In this work, we propose three lattice-free training objectives, namely lattice-free maximum mutual information, lattice-free segment-level minimum Bayes risk, and lattice-free minimum Bayes risk, which are used for the final posterior output of the phoneme-based neural transducer with a limited context dependency. Compared to criteria using N-best lists, lattice-free methods eliminate the decoding step for hypotheses generation during training, which leads to more efficient training. Experimental results show that lattice-free methods gain up to 6.5% relative improvement in word error rate compared to a sequence-level cross-entropy trained model. Compared to the N-best-list based minimum Bayes risk objectives, lattice-free methods gain 40% - 70% relative training time speedup with a small degradation in performance.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ICASSP2
2023 Enhancing and Adversarial: Improve ASR with Speaker Labels
abstract
ASR can be improved by multi-task learning (MTL) with domain enhancing or domain adversarial training, which are two opposite objectives with the aim to increase/decrease domain variance towards domain-aware/agnostic ASR, respectively. In this work, we study how to best apply these two opposite objectives with speaker labels to improve conformer-based ASR. We also propose a novel adaptive gradient reversal layer for stable and effective adversarial training without tuning effort. Detailed analysis and experimental verification are conducted to show the optimal positions in the ASR neural network (NN) to apply speaker enhancing and adversarial training. We also explore their combination for further improvement, achieving the same performance as i-vectors plus adversarial training. Our best speaker-based MTL achieves 7% relative improvement on the Switchboard Hub5’00 set. We also investigate the effect of such speaker-based MTL w.r.t. cleaner dataset and weaker ASR NN.
Wei Zhou 0043, Jingjing Xu 0002, Mohammad Zeineldeen, Christoph Lüscher, Ralf Schlüter, Hermann Ney
ICASSP1
2023 RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech Recognition
abstract
Modern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.For closed-vocabulary scenarios, public tools supporting lexicalconstrained decoding are usually only for classical ASR, or do not support all S2S models.To eliminate this restriction on research possibilities such as modeling unit choice, we present RASR2 in this work, a research-oriented generic S2S decoder implemented in C++.It offers a strong flexibility/compatibility for various S2S models, language models, label units/topologies and neural network architectures.It provides efficient decoding for both open-and closed-vocabulary scenarios based on a generalized search framework with rich support for different search modes and settings.We evaluate RASR2 with a wide range of experiments on both switchboard and Librispeech corpora.Our source code is public online.
Wei Zhou 0043, Eugen Beck, Simon Berger, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2022 On Language Model Integration for RNN Transducer Based Speech Recognition
abstract
The mismatch between an external language model (LM) and the implicitly learned internal LM (ILM) of RNN-Transducer (RNN-T) can limit the performance of LM integration such as simple shallow fusion. A Bayesian interpretation suggests to remove this sequence prior as ILM correction. In this work, we study various ILM correction-based LM integration methods formulated in a common RNN-T framework. We provide a decoding interpretation on two major reasons for performance improvement with ILM correction, which is further experimentally verified with detailed analysis. We also propose an exact-ILM training framework by extending the proof given in the hybrid autoregressive transducer, which enables a theoretical justification for other ILM approaches. Systematic comparison is conducted for both in-domain and cross-domain evaluation on the Librispeech and TED-LIUM Release 2 corpora, respectively. Our proposed exact-ILM training can further improve the best ILM method.
Wei Zhou 0043, Zuoyun Zheng, Ralf Schlüter, Hermann Ney
ICASSP1
2022 Efficient Training of Neural Transducer for Speech Recognition
abstract
As one of the most popular sequence-to-sequence modeling approaches for speech recognition, the RNN-Transducer has achieved evolving performance with more and more sophisticated neural network models of growing size and increasing training epochs. While strong computation resources seem to be the prerequisite of training superior models, we try to overcome it by carefully designing a more efficient training pipeline. In this work, we propose an efficient 3-stage progressive training pipeline to build highly-performing neural transducer models from scratch with very limited computation resources in a reasonable short time period. The effectiveness of each stage is experimentally verified on both Librispeech and Switchboard corpora. The proposed pipeline is able to train transducer models approaching state-of-the-art performance with a single GPU in just 2-3 weeks. Our best conformer transducer achieves 4.1% WER on Librispeech test-other with only 35 epochs of training.
Wei Zhou 0043, Wilfried Michel, Ralf Schlüter, Hermann Ney
INTERSPEECH1
2022 HMM vs. CTC for Automatic Speech Recognition: Comparison Based on Full-Sum Training from Scratch
abstract
In this work, we compare from-scratch sequence-level cross-entropy (full-sum) training of Hidden Markov Model (HMM) and Connectionist Temporal Classification (CTC) topologies for automatic speech recognition (ASR). Besides accuracy, we further analyze their capability for generating high-quality time alignment between the speech signal and the transcription, which can be crucial for many subsequent applications. Moreover, we propose several methods to improve convergence of from-scratch full-sum training by addressing the alignment modeling issue. Systematic comparison is conducted on both Switchboard and LibriSpeech corpora across CTC, posterior HMM with and w/o transition probabilities, and standard hybrid HMM. We also provide a detailed analysis of both Viterbi forced-alignment and Baum-Welch full-sum occupation probabilities.
Tina Raissi, Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney
SLT2
2022 Monotonic Segmental Attention for Automatic Speech Recognition
abstract
We introduce a novel segmental-attention model for automatic speech recognition. We restrict the decoder attention to segments to avoid quadratic runtime of global attention, better generalize to long sequences, and eventually enable streaming. We directly compare global-attention and different segmental-attention modeling variants. We develop and compare two separate time-synchronous decoders, one specifically taking the segmental nature into account, yielding further improvements. Using time-synchronous decoding for segmental models is novel and a step towards streaming applications. Our experiments show the importance of a length model to predict the segment boundaries. The final best segmental-attention model using segmental decoding performs better than global-attention, in contrast to other monotonic attention approaches in the literature. Further, we observe that the segmental model generalizes much better to long sequences of up to several minutes.
Albert Zeyer, Robin Schmitt, Wei Zhou 0043, Ralf Schlüter, Hermann Ney
SLT3
2021 Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
abstract
To join the advantages of classical and end-to-end approaches for speech recognition, we present a simple, novel and competitive approach for phoneme-based neural transducer modeling. Different alignment label topologies are compared and word-end-based phoneme label augmentation is proposed to improve performance. Utilizing the local dependency of phonemes, we adopt a simplified neural network structure and a straightforward integration with the external word-level language model to preserve the consistency of seq-to-seq modeling. We also present a simple, stable and efficient training procedure using frame-wise cross-entropy loss. A phonetic context size of one is shown to be sufficient for the best performance. A simplified scheduled sampling approach is applied for further improvement and different decoding approaches are briefly compared. The overall performance of our best model is comparable to state-of-the-art (SOTA) results for the TED-LIUM Release 2 and Switchboard corpora.
Wei Zhou 0043, Simon Berger, Ralf Schlüter, Hermann Ney
ICASSP1
2021 The Impact of ASR on the Automatic Analysis of Linguistic Complexity and Sophistication in Spontaneous L2 Speech
abstract
In recent years, automated approaches to assessing linguistic complexity in second language (L2) writing have made significant progress in gauging learner performance, predicting human ratings of the quality of learner productions, and benchmarking L2 development.In contrast, there is comparatively little work in the area of speaking, particularly with respect to fully automated approaches to assessing L2 spontaneous speech.While the importance of a well-performing ASR system is widely recognized, little research has been conducted to investigate the impact of its performance on subsequent automatic text analysis.In this paper, we focus on this issue and examine the impact of using a state-of-the-art ASR system for subsequent automatic analysis of linguistic complexity in spontaneously produced L2 speech.A set of 30 selected measures were considered, falling into four categories: syntactic, lexical, n-gram frequency, and information-theoretic measures.The agreement between the scores for these measures obtained on the basis of ASR-generated vs. manual transcriptions was determined through correlation analysis.A more differential effect of ASR performance on specific types of complexity measures when controlling for task type effects is also presented.
Yu Qiao 0005, Wei Zhou 0043, Elma Kerz, Ralf Schlüter
Interspeech2
2021 Equivalence of Segmental and Neural Transducer Modeling: A Proof of Concept
abstract
With the advent of direct models in automatic speech recognition (ASR), the formerly prevalent frame-wise acoustic modeling based on hidden Markov models (HMM) diversified into a number of modeling architectures like encoder-decoder attention models, transducer models and segmental models (direct HMM).While transducer models stay with a frame-level model definition, segmental models are defined on the level of label segments directly.While (soft-)attention-based models avoid explicit alignment, transducer and segmental approach internally do model alignment, either by segment hypotheses or, more implicitly, by emitting so-called blank symbols.In this work, we prove that the widely used class of RNN-Transducer models and segmental models (direct HMM) are equivalent and therefore show equal modeling power.It is shown that blank probabilities translate into segment length probabilities and vice versa.In addition, we provide initial experiments investigating decoding and beam-pruning, comparing time-synchronous and label-/segment-synchronous search strategies and their properties using the same underlying model.
Wei Zhou 0043, Albert Zeyer, André Merboldt, Ralf Schlüter, Hermann Ney
Interspeech1
2021 Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition
abstract
Subword units are commonly used for end-to-end automatic speech recognition (ASR), while a fully acoustic-oriented subword modeling approach is somewhat missing.We propose an acoustic data-driven subword modeling (ADSM) approach that adapts the advantages of several text-based and acousticbased subword methods into one pipeline.With a fully acousticoriented label design and learning process, ADSM produces acoustic-structured subword units and acoustic-matched target sequence for further ASR training.The obtained ADSM labels are evaluated with different end-to-end ASR approaches including CTC, RNN-Transducer and attention models.Experiments on the LibriSpeech corpus show that ADSM clearly outperforms both byte pair encoding (BPE) and pronunciationassisted subword modeling (PASM) in all cases.Detailed analysis shows that ADSM achieves acoustically more logical word segmentation and more balanced sequence length, and thus, is suitable for both time-synchronous and label-synchronous models.We also briefly describe how to apply acoustic-based subword regularization and unseen text segmentation using ADSM.
Wei Zhou 0043, Mohammad Zeineldeen, Zuoyun Zheng, Ralf Schlüter, Hermann Ney
Interspeech1
2020 The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With Specaugment
abstract
We present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is successfully applied to improve performance on top of our best SAT model using i-vectors. By investigating the effect of different maskings, we achieve improvements from SpecAugment on hybrid HMM models without increasing model size and training time. A subsequent sMBR training is applied to fine-tune the final acoustic model, and both LSTM and Transformer language models are trained and evaluated. Our best system achieves a 5.6% WER on the test set, which outperforms the previous state-of-the-art by 27% relative.
Wei Zhou 0043, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlüter, Hermann Ney
ICASSP1
2020 Full-Sum Decoding for Hybrid Hmm Based Speech Recognition Using LSTM Language Model
abstract
In hybrid HMM based speech recognition, LSTM language models have been widely applied and achieved large improvements. The theoretical capability of modeling any unlimited context suggests that no recombination should be applied in decoding. This motivates to reconsider full summation over the HMM-state sequences instead of Viterbi approximation in decoding. We explore the potential gain from more accurate probabilities in terms of decision making and apply the full-sum decoding with a modified prefix-tree search framework. The proposed full-sum decoding is evaluated on both Switchboard and Librispeech corpora. Different models using CE and sMBR training criteria are used. Additionally, both MAP and confusion network decoding as approximated variants of general Bayes decision rule are evaluated. Consistent improvements over strong baselines are achieved in almost all cases without extra cost. We also discuss tuning effort, efficiency and some limitations of full-sum decoding.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
ICASSP1
2020 Robust Beam Search for Encoder-Decoder Attention Based Speech Recognition Without Length Bias
abstract
As one popular modeling approach for end-to-end speech recognition, attention-based encoder-decoder models are known to suffer the length bias and corresponding beam problem.Different approaches have been applied in simple beam search to ease the problem, most of which are heuristic-based and require considerable tuning.We show that heuristics are not proper modeling refinement, which results in severe performance degradation with largely increased beam sizes.We propose a novel beam search derived from reinterpreting the sequence posterior with an explicit length modeling.By applying the reinterpreted probability together with beam pruning, the obtained final probability leads to a robust model modification, which allows reliable comparison among output sequences of different lengths.Experimental verification on the LibriSpeech corpus shows that the proposed approach solves the length bias problem without heuristics or additional tuning effort.It provides robust decision making and consistently good performance under both small and very large beam sizes.Compared with the best results of the heuristic baseline, the proposed approach achieves the same WER on the 'clean' sets and 4% relative improvement on the 'other' sets.We also show that it is more efficient with the additional derived early stopping criterion.
Wei Zhou 0043, Ralf Schlüter, Hermann Ney
INTERSPEECH1