Puming Zhan

dblp:07/4314 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Improving Speed/Accuracy Tradeoff for Online Streaming ASR via Real-Valued and Trainable Strides
abstract
The Conformer Transducer (CT) is arguably the most popular architecture for online streaming end-to-end (E2E) ASR systems. Since it has quadratic complexity in the input sequence length for computing the attention weights, downsampling the input sequence to reduce its length is an effective way to mitigate the computing cost and speed up the inference process. However, in the traditional downsampling approach, the sampling factor (i.e. stride) has to be a pre-defined integer value. The speed up achieved by such kind of downsampling often comes with significant accuracy degradation, because it lacks the flexibility of trading accuracy with speed at fine-grained level. In this paper, we apply the spectral pooling and DiffStride techniques to the CT based online E2E ASR system. This makes the stride a real-valued trainable parameter. We optimize the implementation of these techniques for CT based ASR systems and develop recipes to train the stride together with the model parameters. We conduct experiments on an internal medical conversation dataset. Our results show that we can achieve better tradeoff between recognition accuracy and inference speed by training real-valued stride parameter. Compared to using decimation with integer stride value, our approach reduces real-time factor by 15.6 % on a medical dataset with less than 1 % relative accuracy degradation.
Dario Albesano, Nicola Ferri, Felix Weninger, Puming Zhan
ICASSP4
2022 On the Prediction Network Architecture in RNN-T for ASR
abstract
RNN-T models have gained popularity in the literature and in commercial systems because of their competitiveness and capability of operating in online streaming mode.In this work, we conduct an extensive study comparing several prediction network architectures for both monotonic and original RNN-T models.We compare 4 types of prediction networks based on a common state-of-the-art Conformer encoder and report results obtained on Librispeech and an internal medical conversation data set.Our study covers both offline batch-mode and online streaming scenarios.In contrast to some previous works, our results show that Transformer does not always outperform LSTM when used as prediction network along with Conformer encoder.Inspired by our scoreboard, we propose a new simple prediction network architecture, N -Concat, that outperforms the others in our online streaming benchmark.Transformer and n-gram reduced architectures perform very similarly yet with some important distinct behaviour in terms of previous context.Overall we obtained up to 4.1 % relative WER improvement compared to our LSTM baseline, while reducing prediction network parameters by nearly an order of magnitude (8.4 times).
Dario Albesano, Jesús Andrés-Ferrer, Nicola Ferri, Puming Zhan
INTERSPEECH4
2022 Conformer with dual-mode chunked attention for joint online and offline ASR
abstract
In this paper, we present an in-depth study on online attention mechanisms and distillation techniques for dual-mode (i.e., joint online and offline) ASR using the Conformer Transducer.In the dual-mode Conformer Transducer model, layers can function in online or offline mode while sharing parameters, and in-place knowledge distillation from offline to online mode is applied in training to improve online accuracy.In our study, we first demonstrate accuracy improvements from using chunked attention in the Conformer encoder compared to autoregressive attention with and without lookahead.Furthermore, we explore the efficient KLD and 1-best KLD losses with different shifts between online and offline outputs in the knowledge distillation.Finally, we show that a simplified dual-mode Conformer that only has mode-specific self-attention performs equally well as the one also having mode-specific convolutions and normalization.Our experiments are based on two very different datasets: the Librispeech task and an internal corpus of medical conversations.Results show that the proposed dual-mode system using chunked attention yields 5 % and 4 % relative WER improvement on the Librispeech and medical tasks, compared to the dual-mode system using autoregressive attention with similar average lookahead.
Felix Weninger, Marco Gaudesi, Md. Akmal Haidar, Nicola Ferri, Jesús Andrés-Ferrer, Puming Zhan
INTERSPEECH6
2021 ChannelAugment: Improving Generalization of Multi-Channel ASR by Training with Input Channel Randomization
abstract
End-to-end (E2E) multi-channel ASR systems show state-of-the-art performance in far-field ASR tasks by joint training of a multi-channel front-end along with the ASR model. The main limitation of such systems is that they are usually trained with data from a fixed array geometry, which can lead to degradation in accuracy when a different array is used in testing. This makes it challenging to deploy these systems in practice, as it is costly to retrain and deploy different models for various array configurations. To address this, we present a simple and effective data augmentation technique, which is based on randomly dropping channels in the multi-channel audio input during training, in order to improve the robustness to various array configurations at test time. We call this technique ChannelAugment, in contrast to SpecAugment (SA) which drops time and/or frequency components of a single channel input au-dio. We apply ChannelAugment to the Spatial Filtering (SF) and Minimum Variance Distortionless Response (MVDR) neural beam-forming approaches. For SF, we observe 10.6 % WER improvement across various array configurations employing different numbers of microphones. For MVDR, we achieve a 74 % reduction in training time without causing degradation of recognition accuracy.
Marco Gaudesi, Felix Weninger, Dushyant Sharma, Puming Zhan
ASRU4
2021 Dual-Encoder Architecture with Encoder Selection for Joint Close-Talk and Far-Talk Speech Recognition
abstract
In this paper, we propose a dual-encoder ASR architecture for joint modeling of close-talk (CT) and far-talk (FT) speech, in order to combine the advantages of CT and FT devices for better accuracy. The key idea is to add an encoder selection network to choose the optimal input source (CT or FT) and the corresponding encoder. We use a single-channel encoder for CT speech and a multi-channel encoder with Spatial Filtering neural beamforming for FT speech, which are jointly trained with the encoder selection. We validate our approach on both attention-based and RNN Transducer end-to-end ASR systems. The experiments are done with conversational speech from a medical use case, which is recorded simultaneously with a CT device and a microphone array. Our results show that the proposed dual-encoder architecture obtains up to 9% relative WER reduction when using both CT and FT input, compared to the best single-encoder system trained and tested in matched condition.
Felix Weninger, Marco Gaudesi, Ralf Leibold, Roberto Gemello, Puming Zhan
ASRU5
2021 Contextual Density Ratio for Language Model Biasing of Sequence to Sequence ASR Systems
abstract
End-2-end (E2E) models have become increasingly popular in some ASR tasks because of their performance and advantages. These E2E models directly approximate the posterior distribution of tokens given the acoustic inputs. Consequently, the E2E systems implicitly define a language model (LM) over the output tokens, which makes the exploitation of independently trained language models less straightforward than in conventional ASR systems. This makes it difficult to dynamically adapt E2E ASR system to contextual profiles for better recognizing special words such as named entities. In this work, we propose a contextual density ratio approach for both training a contextual aware E2E model and adapting the language model to named entities. We apply the aforementioned technique to an E2E ASR system, which transcribes doctor and patient conversations, for better adapting the E2E system to the names in the conversations. Our proposed technique achieves a relative improvement of up to 46.5% on the names over an E2E baseline without degrading the overall recognition accuracy of the whole test set. Moreover, it also surpasses a contextual shallow fusion baseline by 22.1 % relative.
Jesús Andrés-Ferrer, Dario Albesano, Puming Zhan, Paul Vozila
Interspeech3
2020 Semi-Supervised Learning with Data Augmentation for End-to-End ASR
abstract
In this paper, we apply Semi-Supervised Learning (SSL) along with Data Augmentation (DA) for improving the accuracy of End-to-End ASR.We focus on the consistency regularization principle, which has been successfully applied to image classification tasks, and present sequence-to-sequence (seq2seq) versions of the FixMatch and Noisy Student algorithms.Specifically, we generate the pseudo labels for the unlabeled data onthe-fly with a seq2seq model after perturbing the input features with DA.We also propose soft label variants of both algorithms to cope with pseudo label errors, showing further performance improvements.We conduct SSL experiments on a conversational speech data set (doctor-patient conversations) with 1.9 kh manually transcribed training data, using only 25 % of the original labels (475 h labeled data).In the result, the Noisy Student algorithm with soft labels and consistency regularization achieves 10.4 % word error rate (WER) reduction when adding 475 h of unlabeled data, corresponding to a recovery rate of 92 %.Furthermore, when iteratively adding 950 h more unlabeled data, our best SSL performance is within 5 % WER increase compared to using the full labeled training set (recovery rate: 78 %).
Felix Weninger, Franco Mana, Roberto Gemello, Jesús Andrés-Ferrer, Puming Zhan
INTERSPEECH5
2019 Online Batch Normalization Adaptation for Automatic Speech Recognition
abstract
Deep Neural Network (DNN) acoustic models are sensitive to the mismatch between training and testing environments. When a trained model is tested on unseen speakers, domain, or environment, recognition accuracy can degrade substantially. In such a case, offline adaptation with a fair amount of field data can improve recognition accuracy significantly, and is commonly applied to ASR systems in practice. Ideally, such kind of adaptation should be done online as well in order to catch any unexpected dynamic changes in the environments during the inference process. However, online adaptation is subject to strict constraints on computational cost. On the other hand, the small amount of available data and the nature of unsupervised adaptation make online adaptation a very challenging task, especially for DNN acoustic models which normally contain millions of parameters. In this paper, we introduce a simple and effective online adaptation technique to compensate training and testing mismatch for DNN acoustic models. It is done via online adaptation of the parameters associated with the batch normalization applied to the model training process. Our results show that this technique can improve accuracy significantly in a domain mismatched scenario for different DNN architectures.
Franco Mana, Felix Weninger, Roberto Gemello, Puming Zhan
ASRU4
2019 Listen, Attend, Spell and Adapt: Speaker Adapted Sequence-to-Sequence ASR
abstract
Sequence-to-sequence (seq2seq) based ASR systems have shown state-of-the-art performances while having clear advantages in terms of simplicity.However, comparisons are mostly done on speaker independent (SI) ASR systems, though speaker adapted conventional systems are commonly used in practice for improving robustness to speaker and environment variations.In this paper, we apply speaker adaptation to seq2seq models with the goal of matching the performance of conventional ASR adaptation.Specifically, we investigate Kullback-Leibler divergence (KLD) as well as Linear Hidden Network (LHN) based adaptation for seq2seq ASR, using different amounts (up to 20 hours) of adaptation data per speaker.Our SI models are trained on large amounts of dictation data and achieve state-of-the-art results.We obtained 25% relative word error rate (WER) improvement with KLD adaptation of the seq2seq model vs. 18.7% gain from acoustic model adaptation in the conventional system.We also show that the WER of the seq2seq model decreases log-linearly with the amount of adaptation data.Finally, we analyze adaptation based on the minimum WER criterion and adapting the language model (LM) for score fusion with the speaker adapted seq2seq model, which result in further improvements of the seq2seq system performance.
Felix Weninger, Jesús Andrés-Ferrer, Puming Zhan
INTERSPEECH4
2019 Deep Learning Based Mandarin Accent Identification for Accent Robust ASR
Felix Weninger, Daniel Willett, Puming Zhan
INTERSPEECH5
2017 Semi-supervised training strategies for deep neural networks
abstract
Use of both manually and automatically labelled data for model training is referred to as semi-supervised training. While semi-supervised acoustic model training has been well-explored in the context of hidden Markov Gaussian mixture models (HMM-GMMs), the re-emergence of deep neural network (DNN) acoustic models has given rise to some novel approaches to semi-supervised DNN training. This paper investigates several different strategies for semi-supervised DNN training, including the so-called `shared hidden layer' approach and the `knowledge distillation' (or student-teacher) approach. Particular attention is paid to the differing behaviour of semi-supervised DNN training methods during the cross-entropy and sequence training phases of model building. Experimental results on our internal study dataset provide evidence that in a low-resource scenario the most effective semi-supervised training strategy is `naive CE' (treating manually transcribed and automatically transcribed data identically during the cross entropy phase of training) followed by use of a shared hidden layer technique during sequence training.
Matthew Gibson, Gary Cook, Puming Zhan
ASRU3
2012 Constructing ensembles of dissimilar acoustic models using hidden attributes of training data
abstract
One of the objectives in acoustic modeling is to realize robust statistical models against the wide variety of acoustic conditions that are present in real world environments. As large amounts of training data become available, modeling subsets of the data with similar acoustic qualities can be done accurately and multiple acoustic models are jointly used as a form of system combination or model selection. In this paper, we propose a method to partition the training data for constructing ensembles of acoustic models using metadata attributes such as SNR, speaking rate, and duration via a binary tree. The metadata attribute used at each binary split in the decision tree is obtained using a metric proposed in this paper that is cosine-similarity based. The resulting multiple models are combined using voting techniques such as n-best ROVER. The proposed method improved the recognition accuracy by up to 4% relative over the state-of-the-art system on a large vocabulary continuous speech recognition voice search task.
Takashi Fukuda, Ryuki Tachibana, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan
ICASSP5
2011 Frame-level AnyBoost for LVCSR with the MMI Criterion
abstract
This paper propose a variant of AnyBoost for a large vocabulary continuous speech recognition (LVCSR) task. AnyBoost is an efficient algorithm to train an ensemble of weak learners by gradient descent for an objective function.We present a novel training procedure that trains acoustic models via the MMI criterion using data that is weighted proportional to the summation of the posterior functions of previous round of weak learners. Optimized for system combination by n-best ROVER at runtime, data weights for a new weak learner are computed as a weighted summation of posteriors of previous weak learners. We compare a frame-based version and a sentence-based version of our proposed algorithm with a frame-based AdaBoost algorithm. We will present results on a voice search task trained with different amounts of data with gains of 5.1% to 7.5% relative in WER can be obtained by three rounds of boosting.
Ryuki Tachibana, Takashi Fukuda, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan
ASRU5
1999 Progress in Broadcast News transcription at Dragon Systems
abstract
We report on progress in acoustic modelling and preprocessing in our Broadcast News transcription system. We have gone back to basics in acoustic modelling, and re-examined some of our standard practices, in particular the use of IMELDA and frequency warping, in the context of the Broadcast News corpus. We also report on some preliminary experiments with a generalization of IMELDA, "semi-tied covariances". In combination, these improvements lead to a 3.5% absolute improvement over our eval97 models. We also describe our attempts to fix our rather primitive, silence-based preprocessing system, including initial results using a new speaker-change detection algorithm based on Hotelling's T/sup 2/-test.
Steven Wegmann, Puming Zhan, Larry Gillick
ICASSP2
1999 Dragon systems' 1998 broadcast news transcription system
abstract
In this paper we shall describe key improvements to Dragon’s Broadcast News Transcription System, which include: the addition of a speaker-change detection algorithm to our preprocessing subsystem, a new diagonalizing transformation trained using semi-tied covariances, and the addition of probabilities on pronunciations. This new transcription system yields a word error rate of 15.2% on the 1997 evaluation test data, and 14.5% őn the 1998 evaluation test data.
Steven Wegmann, Puming Zhan, Ira Carp, Michael Newman, Jonathan Yamron, Larry Gillick
EUROSPEECH2
1997 Janus-III: speech-to-speech translation in multiple languages
abstract
This paper describes JANUS-III, our most recent version of the JANUS speech-to-speech translation system. We present an overview of the system and focus on how system design facilitates speech translation between multiple languages, and allows for easy adaptation to new source and target languages. We also describe our methodology for evaluation of end-to-end system performance with a variety of source and target languages. For system development and evaluation, we have experimented with both push-to-talk as well as cross-talk recording conditions. To date, our system has achieved performance levels of over 80% acceptable translations on transcribed input, and over 70% acceptable translations on speech input recognized with a 75-90% word accuracy. Our current major research is concentrated on enhancing the capabilities of the system to deal with input in broad and general domains.
Alon Lavie, Alex Waibel, Lori S. Levin, Michael Finke, Donna Gates, Marsal Gavaldà, Torsten Zeppenfeld, Puming Zhan
ICASSP8
1997 Speaker normalization based on frequency warping
abstract
In speech recognition, speaker-dependence of a speech recognition system comes from speaker-dependence of the speech feature, and the variation of vocal tract shape is the major source of inter-speaker variations of the speech feature, though there are some other sources which also contribute. In this paper, we address the approach of speaker normalization which aims at normalizing speaker's vocal tract length based on frequency warping (FWP). The FWP is implemented in the front-end preprocessing of our speech recognition system. We investigate the formant-based and ML-based FWP in linear and nonlinear warping modes, and compare them in detail. All experimental results are based on our JANUS3 large vocabulary continuous speech recognition system and the Spanish Spontaneous Scheduling Task database (SSST).
Puming Zhan, Martin Westphal
ICASSP1
1997 Speaker normalization and speaker adaptation - a combination for conversational speech recognition
abstract
Speaker normalization and speaker adaptation are two strategies to tackle the variations from speaker, channel, and environment. The vocal tract length normalization (VTLN) is an effective speaker normalization approach to compensate for the variations of vocal tract shapes. The Maximum Likelihood Linear Regression(MLLR) is a recent proposed method for speaker-adaptation. In this paper, we propose a speaker-specific Bark scale VTLN method, investigate the combination of the VTLN with MLLR, and present an iterative procedure for decoding the combined system of VTLN and MLLR. The results show that: (1) the new VTLN method is very effective with which the word error rate can be reduced up to 11%; (2) the combination of VTLN and MLLR can provide up to 15% word error reduction; (3) both VTLN and MLLR are more effective for the push-to-talk data than for the cross-talk data. 1 INTRODUCTION Almost all speech recognizers are, in some extent, sensitive to the variations of speakers and/or env...
Puming Zhan, Martin Westphal, Michael Finke, Alex Waibel
EUROSPEECH1
1996 JANUS-II-translation of spontaneous conversational speech
abstract
JANUS-II is a research system to design and test components of speech-to-speech translation systems as well as a research prototype for such a system. We focus on two aspects of the system: (1) the new features of the speech recognition component JANUS-SR, and (2) the end-to-end performance of JANUS-II, including a comparison of two machine translation strategies used for JANUS-MT (PHOENIX and GLR*).
Alex Waibel, Michael Finke, Donna Gates, Marsal Gavaldà, Thomas Kemp, Alon Lavie, Lori S. Levin, Laura Mayfield Tomokiyo, Arthur E. McNair, Ivica Rogina, Kaori Shima, Tilo Sloboda, Monika Woszczyna, Torsten Zeppenfeld, Puming Zhan
ICASSP16
1996 Translation of conversational speech with JANUS-II
abstract
In this paper we investigate the possibility of translating continuous spoken conversations in a cross-talk environment.This is a task known to be difficult for human translators due to several factors.It is characterized by rapid and even overlapping turn-taking, a high degree of co-articulation, and fragmentary language.We describe experiments using both push-to-talk as well as cross-talk recording conditions.Our results indicate that conversational speech recognition and translation is possible, even in a free crosstalk environment.To date, our system has achieved performances of over 80% acceptable translations on transcribed input, and over 70% acceptable translations on speech input recognized with a 70-80% word accuracy.The system's performance on spontaneous conversations recorded in a cross-talk environment is shown to be as good and even slightly superior to the simpler and easier push-to-talk scenario.
Alon Lavie, Alex Waibel, Lori S. Levin, Donna Gates, Marsal Gavaldà, Torsten Zeppenfeld, Puming Zhan, Oren Glickman
ICSLP7
1996 JANUS-II: towards spontaneous Spanish speech recognition
abstract
JANUS-II is a research system for investigating various issues in speech-to-speech translations and has been implemented for speech-to-speech translations on many languages 1 .In this paper, we address the Spanish speech recognition part of JANUS-II.First, we report the bootstrap and optimization of the recognition system.Then we i n v estigate the di erence between push-to-talk and cross-talk dialogs, which are two di erent kinds of data in our database.We give a detail noise analysis for the push-to-talk and cross-talk dialogs and present some recognition results for the comparison.We h a v e observed that the cross-talk dialogs are harder than the pushto-talk dialogs for speech recognition, because they are more noisy than the latter.Currently, the error rate of our Spanish recognizer is 27 for push-to-talk test set and 32 for crosstalk test set.
Puming Zhan, Klaus Ries 0001, Marsal Gavaldà, Donna Gates, Alon Lavie, Alex Waibel
ICSLP1