VLDB 2026 Research / reviewers in the wild / expert
Zoltán Tüske
dblp:79/4933
· DBLP profile ↗
53ranked-venue papers
20as first author
10since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 20 first-author · 10 since 2021Artificial intelligence and machine learning · 36 · 12 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SADA: Saudi Audio Dataset for ArabicabstractArabic is among the most challenging languages in the world. Unfortunately, the scarcity of Arabic datasets makes studies in Arabic speech technology demanding. This paper introduces SADA, the Saudi Audio Dataset for Arabic, with 668 hours of high-quality audio suitable for supervised training. The audio recordings were sourced from 57 television shows provided by the Saudi Broadcasting Authority. The audio covers both read and spontaneous speaking styles in various genres. The National Center for Artificial Intelligence in Saudi Arabia transcribed and prepared the data for training and processing. The recordings are in Arabic. Most are in Saudi dialects, while other Arabic dialects include Yemeni, Egyptian, and Levantine. The dataset is split into training, validation, and testing sets to enhance its usage. The validation and testing sets contain 10 hours of audio segments each. Besides giving a detailed description of the dataset, wide range of speech recognition experiments using standard tools are also presented. Sadeen Alharbi, Areeb Alowisheq, Zoltán Tüske, Kareem Darwish, Abdullah Alrajeh, Abdulmajeed Alrowithi, Aljawharah Bin Tamran, Asma Ibrahim, Raghad Aloraini, Raneem Alnajim, Ranya A. Alkahtani, Renad Almuasaad, Sara Alrasheed, Shaykhah Alsubaie, Yaser Alonaizan |
ICASSP | 3 |
| 2022 | Improving End-to-end Models for Set Prediction in Spoken Language UnderstandingabstractThe goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far cheaper to collect than verbatim transcripts. We focus on this set prediction problem, where entity order is unspecified. Using two classes of E2E models, RNN transducers and attention based encoder-decoders, we show that these models work best when the training entity sequence is arranged in spoken order. To improve E2E SLU models when entity spoken order is unknown, we propose a novel data augmentation technique along with an implicit attention based alignment method to infer the spoken order. F1 scores significantly increased by more than 11% for RNN-T and about 2% for attention based encoder-decoder SLU models, outperforming previously reported results. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Brian Kingsbury, George Saon |
ICASSP | 2 |
| 2021 | RNN Transducer Models for Spoken Language UnderstandingabstractWe present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
ICASSP | 4 |
| 2021 | End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained FeaturesabstractTransformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of spoken language understanding (SLU) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SLU transformer network based architecture which allows the use of self-supervised pre- trained acoustic features, pre-trained model initialization and multi-task training. Several SLU experiments for predicting intent and entity labels/values using the ATIS dataset are performed. These experiments investigate the interaction of pre-trained model initialization and multi-task training with either traditional filterbank or self-supervised pre-trained acoustic features. Results show not only that self-supervised pre-trained acoustic features outperform filterbank features in almost all the experiments, but also that when these features are used in combination with multi-task training, they almost eliminate the necessity of pre-trained model initialization. Edmilson da Silva Morais, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zoltán Tüske, Brian Kingsbury |
ICASSP | 4 |
| 2021 | Advancing RNN Transducer Technology for Speech RecognitionabstractWe investigate a set of techniques for RNN Transducers (RNN-Ts) that were instrumental in lowering the word error rate on three different tasks (Switchboard 300 hours, conversational Spanish 780 hours and conversational Italian 900 hours). The techniques pertain to architectural changes, speaker adaptation, language model fusion, model combination and general training recipe. First, we introduce a novel multiplicative integration of the encoder and prediction network vectors in the joint network (as opposed to additive). Second, we discuss the applicability of i-vector speaker adaptation to RNN-Ts in conjunction with data perturbation. Third, we explore the effectiveness of the recently proposed density ratio language model fusion for these tasks. Last but not least, we describe the other components of our training recipe and their effect on recognition performance. We report a 5.9% and 12.5% word error rate on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation and a 12.7% WER on the Mozilla CommonVoice Italian test set. George Saon, Zoltán Tüske, Daniel Bolaños, Brian Kingsbury |
ICASSP | 2 |
| 2021 | Reducing Exposure Bias in Training Recurrent Neural Network TransducersabstractWhen recurrent neural network transducers (RNNTs) are trained using the typical maximum likelihood criterion, the prediction network is trained only on ground truth label sequences.This leads to a mismatch during inference, known as exposure bias, when the model must deal with label sequences containing errors.In this paper we investigate approaches to reducing exposure bias in training to improve the generalization of RNNT models for automatic speech recognition (ASR).A label-preserving input perturbation to the prediction network is introduced.The input token sequences are perturbed using SwitchOut and scheduled sampling based on an additional token language model.Experiments conducted on the 300-hour Switchboard dataset demonstrate their effectiveness.By reducing the exposure bias, we show that we can further improve the accuracy of a high-performance RNNT ASR model and obtain state-of-the-art results on the 300-hour Switchboard dataset. Brian Kingsbury, George Saon, David Haws, Zoltán Tüske |
Interspeech | 5 |
| 2021 | 4-Bit Quantization of LSTM-Based Speech Recognition ModelsabstractWe investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition (ASR): hybrid Deep Bidirectional LSTM -Hidden Markov Models (DBLSTM-HMMs) and Recurrent Neural Network -Transducers (RNN-Ts).Using a 4-bit integer representation, a naïve quantization approach applied to the LSTM portion of these models results in significant Word Error Rate (WER) degradation.On the other hand, we show that minimal accuracy loss is achievable with an appropriate choice of quantizers and initializations.In particular, we customize quantization schemes depending on the local properties of the network, improving recognition performance while limiting computational time.We demonstrate our solution on the Switchboard (SWB) and CallHome (CH) test sets of the NIST Hub5-2000 evaluation.DBLSTM-HMMs trained with 300 or 2000 hours of SWB data achieves <0.5% and <1% average WER degradation, respectively.On the more challenging RNN-T models, our quantization strategy limits degradation in 4-bit inference to 1.3%. Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Xiao Sun 0013, Naigang Wang, Swagath Venkataramani, George Saon, Brian Kingsbury, Wei Zhang 0022, Zoltán Tüske, Kailash Gopalakrishnan |
Interspeech | 11 |
| 2021 | Integrating Dialog History into End-to-End Spoken Language Understanding SystemsabstractEnd-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation independently. Spoken conversations on the other hand, are very much context dependent, and dialog history contains useful information that can improve the processing of each conversational turn. In this paper, we investigate the importance of dialog history and how it can be effectively integrated into end-to-end SLU systems. While processing a spoken utterance, our proposed RNN transducer (RNN-T) based SLU model has access to its dialog history in the form of decoded transcripts and SLU labels of previous turns. We encode the dialog history as BERT embeddings, and use them as an additional input to the SLU model along with the speech features for the current utterance. We evaluate our approach on a recently released spoken dialog data set, the HarperValleyBank corpus. We observe significant improvements: 8% for dialog action and 30% for caller intent recognition tasks, in comparison to a competitive context independent end-to-end baseline system. Jatin Ganhotra, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltán Tüske, Brian Kingsbury |
Interspeech | 6 |
| 2021 | Improving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio
Gakuto Kurata, George Saon, Brian Kingsbury, David Haws, Zoltán Tüske |
Interspeech | 5 |
| 2021 | On the Limit of English Conversational Speech RecognitionabstractIn our previous work we demonstrated that a single headed attention encoder-decoder model is able to reach state-of-the-art results in conversational speech recognition. In this paper, we further improve the results for both Switchboard 300 and 2000. Through use of an improved optimizer, speaker vector embeddings, and alternative speech representations we reduce the recognition errors of our LSTM system on Switchboard-300 by 4% relative. Compensation of the decoder model with the probability ratio approach allows more efficient integration of an external language model, and we report 5.9% and 11.5% WER on the SWB and CHM parts of Hub5'00 with very simple LSTM models. Our study also considers the recently proposed conformer, and more advanced self-attention based language models. Overall, the conformer shows similar performance to the LSTM; nevertheless, their combination and decoding with an improved LM reaches a new record on Switchboard-300, 5.0% and 10.0% WER on SWB and CHM. Our findings are also confirmed on Switchboard-2000, and a new state of the art is reported, practically reaching the limit of the benchmark. Zoltán Tüske, George Saon, Brian Kingsbury |
Interspeech | 1 |
| 2020 | Alignment-Length Synchronous Decoding for RNN TransducerabstractWe present a beam decoding strategy for recurrent neural network transducers which has the characteristic that all competing hypotheses within the beam have the same alignment length (number of output symbols plus BLANK symbols). We contrast the proposed technique with time-synchronous decoding where the competing hypotheses within the beam correspond to the same input frames (but can have different length output sequences). Experiments on the Switchboard 2000 hours corpus show that alignment-length synchronous decoding (ALSD) is 25% faster than time-synchronous decoding (TSD) for the same accuracy because ALSD performs 42% fewer joint network evaluations and hypothesis expansions during the search. Additionally, we discuss the benefit of caching and batching the prediction and joint network evaluations, of using prefix trees instead of full output vocabulary expansions, and of performing hypothesis recombination after pruning. With open beam decoding, we reach a 6.2% / 10.9% word error rate on the Switchboard and CallHome Hub5 2000 evaluation testsets which compares favorably to other published single-model results on this corpus. George Saon, Zoltán Tüske, Kartik Audhkhasi |
ICASSP | 2 |
| 2020 | End-to-End Spoken Language Understanding Without Full TranscriptsabstractAn essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
INTERSPEECH | 2 |
| 2020 | Single Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on SwitchboardabstractIt is generally believed that direct sequence-to-sequence (seq2seq) speech recognition models are competitive with hybrid models only when a large amount of data, at least a thousand hours, is available for training. In this paper, we show that state-of-the-art recognition performance can be achieved on the Switchboard-300 database using a single headed attention, LSTM based model. Using a cross-utterance language model, our single-pass speaker independent system reaches 6.4% and 12.5% word error rate (WER) on the Switchboard and CallHome subsets of Hub5'00, without a pronunciation lexicon. While careful regularization and data augmentation are crucial in achieving this level of performance, experiments on Switchboard-2000 show that nothing is more useful than more data. Overall, the combination of various regularizations and a simple but fairly large model results in a new state of the art, 4.7% and 7.8% WER on the Switchboard and CallHome sets, using SWB-2000 without any external data resources. Zoltán Tüske, George Saon, Kartik Audhkhasi, Brian Kingsbury |
INTERSPEECH | 1 |
| 2019 | Semi-Supervised Training and Data Augmentation for Adaptation of Automatic Broadcast News Captioning SystemsabstractIn this paper we present a comprehensive study on building and adapting deep neural network based speech recognition systems for automatic closed captioning. We develop the proposed systems by first building base automatic speech recognition (ASR) systems that are not specific to any particular show or station. These models are trained on nearly 6000 hours of broadcast news data using conventional hybrid and more recent attention based end-to-end acoustic models. We then employ various adaptation and data augmentation strategies to further improve the trained base models. We use 535 hours of data from two independent BN sources to study how the base models can be customized. We observe up to 32% relative improvement using the proposed techniques on test sets related to, but independent of the adaptation data. At these low word error rates (WERs), we believe the customized BN ASR systems can be used effectively for automatic closed captioning. Samuel Thomas 0001, Masayuki Suzuki, Zoltán Tüske, Larry Sansone, Michael Picheny |
ASRU | 4 |
| 2019 | Simplified LSTMS for Speech RecognitionabstractIn this paper we explore new variants of Long Short-Term Memory (LSTM) networks for sequential modeling of acoustic features. In particular, we show that: (i) removing the output gate, (ii) replacing the hyperbolic tangent nonlinearity at the cell output with hard tanh, and (iii) collapsing the cell and hidden state vectors leads to a model that is conceptually simpler than and comparable in effectiveness to a regular LSTM for speech recognition. The proposed model has 25% fewer parameters than an LSTM with the same number of cells, trains faster because it has larger gradients leading to larger steps in weight space, and reaches a better optimum because there are fewer nonlinearities to traverse across layers. We report experimental results for both hybrid and CTC acoustic models on three publicly available English datasets: Switchboard 300 hours telephone conversations, 400 hours broadcast news transcription, and the MALACH 176 hours corpus of Holocaust survivor testimonies. In all cases the proposed models achieve similar or better accuracy than regular LSTMs while being conceptually simpler. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas 0001 |
ASRU | 2 |
| 2019 | Sequence Noise Injected Training for End-to-end Speech RecognitionabstractWe present a simple noise injection algorithm for training end-to-end ASR models which consists in adding to the spectra of training utterances the scaled spectra of random utterances of comparable length. We conjecture that the sequence information of the "noise" utterances is important and verify this via a contrast experiment where the frames of the utterances to be added are randomly shuffled. Experiments for both CTC and attention-based models show that the pro-posed scheme results in up to 9% relative word error rate improvements (depending on the model and test set) on the Switchboard 300 hours English conversational telephony database. Additionally, we set a new benchmark for attention-based encoder-decoder models on this corpus. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury |
ICASSP | 2 |
| 2019 | English Broadcast News Speech Recognition by Humans and MachinesabstractWith recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broadcast news (BN), a similar challenging task. We also perform a set of recognition measurements to understand how close the achieved automatic speech recognition results are to human performance on this task. On two publicly available BN test sets, DEV04F and RT04, our speech recognition system using LSTM and residual network based acoustic models with a combination of n-gram and neural network language models performs at 6.5% and 5.9% word error rate. By achieving new performance milestones on these test sets, our experiments show that techniques developed on other related tasks, like CTS, can be transferred to achieve similar performance. In contrast, the best measured human recognition performance on these test sets is much lower, at 3.6% and 2.8% respectively, indicating that there is still room for new techniques and improvements in this space, to reach human performance levels. Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, Zoltán Tüske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
ICASSP | 5 |
| 2019 | Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition
Kartik Audhkhasi, George Saon, Zoltán Tüske, Brian Kingsbury, Michael Picheny |
INTERSPEECH | 3 |
| 2019 | Challenging the Boundaries of Speech Recognition: The MALACH CorpusabstractThere has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour subset of a large archive of Holocaust testimonies collected by the Survivors of the Shoah Visual History Foundation, presents significant challenges to the speech community. The collection consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching, and emotional speech - all still open problems for speech recognition systems. Transcription is challenging even for skilled human annotators. This paper proposes that the community place focus on the MALACH corpus to develop speech recognition systems that are more robust with respect to accents, disfluencies and emotional speech. To reduce the barrier for entry, a lexicon and training and testing setups have been created and baseline results using current deep learning technologies are presented. The metadata has just been released by LDC (LDC2019S11). It is hoped that this resource will enable the community to build on top of these baselines so that the extremely important information in these and related oral histories becomes accessible to a wider audience. Michael Picheny, Zoltán Tüske, Brian Kingsbury, Kartik Audhkhasi, George Saon |
INTERSPEECH | 2 |
| 2019 | Detection and Recovery of OOVs for Improved English Broadcast News Captioning
Samuel Thomas 0001, Kartik Audhkhasi, Zoltán Tüske, Michael Picheny |
INTERSPEECH | 3 |
| 2019 | Advancing Sequence-to-Sequence Based Speech Recognition
Zoltán Tüske, Kartik Audhkhasi, George Saon |
INTERSPEECH | 1 |
| 2018 | Acoustic Modeling of Speech Waveform Based on Multi-Resolution, Neural Network Signal ProcessingabstractRecently, several papers have demonstrated that neural networks (NN) are able to perform the feature extraction as part of the acoustic model. Motivated by the Gammatone feature extraction pipeline, in this paper we extend the waveform based NN model by a second level of time-convolutional element. The proposed extension generalizes the envelope extraction block, and allows the model to learn multi-resolutional representations. Automatic speech recognition (ASR) experiments show significant word error rate reduction over our previous best acoustic model trained in the signal domain directly. Although we use only 250 hours of speech, the data-driven NN based speech signal processing performs nearly equally to traditional handcrafted feature extractors. In additional experiments, we also test segment-level feature normalization techniques on NN derived features, which improve the results further. However, the porting of speech representations derived by a feed-forward NN to a LSTM back-end model indicates much less robustness of the NN front-end compared to the standard feature extractors. Analysis of the weights in the proposed new layer reveals that the NN prefers both multi-resolution and modulation spectrum representations. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2018 | Investigation on LSTM Recurrent N-gram Language Models for Speech RecognitionabstractRecurrent neural networks (NN) with long short-term memory (LSTM) are the current state of the art to model long term dependencies.However, recent studies indicate that NN language models (LM) need only limited length of history to achieve excellent performance.In this paper, we extend the previous investigation on LSTM network based n-gram modeling to the domain of automatic speech recognition (ASR).First, applying recent optimization techniques and up to 6-layer LSTM networks, we improve LM perplexities by nearly 50% relative compared to classic count models on three different domains.Then, we demonstrate by experimental results that perplexities improve significantly only up to 40-grams when limiting the LM history.Nevertheless, the ASR performance saturates already around 20-grams despite across sentence modeling.Analysis indicates that the performance gain of LSTM NNLM over count models results only partially from the longer context and cross sentence modeling capabilities.Using equal context, we show that deep 4-gram LSTM can significantly outperform large interpolated count models by performing the backing off and smoothing significantly better.This observation also underlines the decreasing importance to combine state-of-the-art deep NNLM with count based model. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2018 | Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural NetworksabstractWe propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolution step to improve the average fitness of the population. With a back-off strategy in the SGD step and an elitist strategy in the evolution step, it guarantees that the best fitness in the population will never degrade. In addition, individuals in the population optimized with various SGD-based optimizers using distinct hyper-parameters in the SGD step are considered as competing species in a coevolution setting such that the complementarity of the optimizers is also taken into account. The effectiveness of ESGD is demonstrated across multiple applications including speech recognition, image recognition and language modeling, using networks with a variety of deep architectures. Wei Zhang 0022, Zoltán Tüske, Michael Picheny |
NeurIPS | 3 |
| 2017 | Parallel Neural Network Features for Improved Tandem Acoustic ModelingabstractThe combination of acoustic models or features is a standard approach to exploit various knowledge sources.This paper investigates the concatenation of different bottleneck (BN) neural network (NN) outputs for tandem acoustic modeling.Thus, combination of NN features is performed via Gaussian mixture models (GMM).Complementarity between the NN feature representations is attained by using various network topologies: LSTM recurrent, feed-forward, and hierarchical, as well as different non-linearities: hyperbolic tangent, sigmoid, and rectified linear units.Speech recognition experiments are carried out on various tasks: telephone conversations, Skype calls, as well as broadcast news and conversations.Results indicate that LSTM based tandem approach is still competitive, and such tandem model can challenge comparable hybrid systems.The traditional steps of tandem modeling, speaker adaptive and sequence discriminative GMM training, improve the tandem results further.Furthermore, these "old-fashioned" steps remain applicable after the concatenation of multiple neural network feature streams.Exploiting the parallel processing of input feature streams, it is shown that 2-5% relative improvement could be achieved over the single best BN feature set.Finally, we also report results after neural network based language model rescoring and examine the system combination possibilities using such complex tandem models. Zoltán Tüske, Wilfried Michel, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2016 | Investigation on log-linear interpolation of multi-domain neural network language modelabstractInspired by the success of multi-task training in acoustic modeling, this paper investigates a new architecture for a multi-domain neural network based language model (NNLM). The proposed model has several shared hidden layers and domain-specific output layers. As will be shown, the log-linear interpolation of the multi-domain outputs and the optimization of interpolation weights fit naturally in the framework of NNLM. The resulting model can be expressed as a single NNLM. As an initial study of such an architecture, this paper focuses on deep feed-forward neural networks (DNNs). We also re-investigate the potential of long context up to 30-grams, and depth up to 5 hidden layers in DNN-LM. Our final feed-forward multidomain NNLM is trained on 3.1B running words across 11 domains for English broadcast news and conversations large vocabulary continuous speech recognition task. After log-linear interpolation and fine-tuning, we measured improvements in terms of perplexity and word error rate over the models trained on 50M running words of in-domain news resources. The final multi-domain feed-forward LM outperformed our previous best LSTM-RNN LM trained on the 50M in-domain corpus, even after linear interpolation with large count models. Zoltán Tüske, Kazuki Irie, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2016 | LSTM, GRU, Highway and a Bit of Attention: An Empirical Overview for Language Modeling in Speech RecognitionabstractPopularized by the long short-term memory (LSTM), multiplicative gates have become a standard means to design artificial neural networks with intentionally organized information flow.Notable examples of such architectures include gated recurrent units (GRU) and highway networks.In this work, we first focus on the evaluation of each of the classical gated architectures for language modeling for large vocabulary speech recognition.Namely, we evaluate the highway network, lateral network, LSTM and GRU.Furthermore, the motivation underlying the highway network also applies to LSTM and GRU.An extension specific to the LSTM has been recently proposed with an additional highway connection between the memory cells of adjacent LSTM layers.In contrast, we investigate an approach which can be used with both LSTM and GRU: a highway network in which the LSTM or GRU is used as the transformation function.We found that the highway connections enable both standalone feedforward and recurrent neural language models to benefit better from the deep structure and provide a slight improvement of recognition accuracy after interpolation with count models.To complete the overview, we include our initial investigations on the use of the attention mechanism for learning word triggers. Kazuki Irie, Zoltán Tüske, Tamer Alkhouli, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2015 | Multilingual representations for low resource speech recognition and keyword searchabstractThis paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland |
ASRU | 11 |
| 2015 | Speaker adaptive joint training of Gaussian mixture models and bottleneck featuresabstractIn the tandem approach, the output of a neural network (NN) serves as input features to a Gaussian mixture model (GMM) aiming to improve the emission probability estimates. As has been shown in our previous work, GMM with pooled covariance matrix can be integrated into a neural network framework as a softmax layer with hidden variables, which allows for joint estimation of both neural network and Gaussian mixture parameters. Here, this approach is extended to include speaker adaptive training (SAT) by introducing a speaker dependent neural network layer. Error backpropagation beyond this speaker dependent layer realizes the adaptive training of the Gaussian parameters as well as the optimization of the bottleneck (BN) tandem features of the underlying acoustic model, simultaneously. In this study, after the initialization by constrained maximum likelihood linear regression (CMLLR) the speaker dependent layer itself is kept constant during the joint training. Experiments show that the deeper backpropagation through the speaker dependent layer is necessary for improved recognition performance. The speaker adaptively and jointly trained BN-GMM results in 5% relative improvement over very strong speaker-independent hybrid baseline on the Quaero English broadcast news and conversations task, and on the 300-hour Switchboard task. Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney |
ASRU | 1 |
| 2015 | Integrating Gaussian mixtures into deep neural networks: Softmax layer with hidden variablesabstractIn the hybrid approach, neural network output directly serves as hidden Markov model (HMM) state posterior probability estimates. In contrast to this, in the tandem approach neural network output is used as input features to improve classic Gaussian mixture model (GMM) based emission probability estimates. This paper shows that GMM can be easily integrated into the deep neural network framework. By exploiting its equivalence with the log-linear mixture model (LMM), GMM can be transformed to a large softmax layer followed by a summation pooling layer. Theoretical and experimental results indicate that the jointly trained and optimally chosen GMM and bottleneck tandem features cannot perform worse than a hybrid model. Thus, the question “hybrid vs. tandem” simplifies to optimizing the output layer of a neural network. Speech recognition experiments are carried out on a broadcast news and conversations task using up to 12 feed-forward hidden layers with sigmoid and rectified linear unit activation functions. The evaluation of the LMM layer shows recognition gains over the classic softmax output. Zoltán Tüske, Muhammad Ali Tahir, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2015 | Convolutional neural networks for acoustic modeling of raw time signal in LVCSRabstractIn this paper we continue to investigate how the deep neural network (DNN) based acoustic models for automatic speech recognition can be trained without hand-crafted feature extraction. Previously, we have shown that a simple fully connected feedforward DNN performs surprisingly well when trained directly on the raw time signal. The analysis of the weights revealed that the DNN has learned a kind of short-time time-frequency decomposition of the speech signal. In conventional feature extraction pipelines this is done manually by means of a filter bank that is shared between the neighboring analysis windows. Following this idea, we show that the performance gap between DNNs trained on spliced hand-crafted features and DNNs trained on raw time signal can be strongly reduced by introducing 1D-convolutional layers. Thus, the DNN is forced to learn a short-time filter bank shared over a longer time span. This also allows us to interpret the weights of the second convolutional layer in the same way as 2D patches learned on critical band energies by typical convolutional neural networks. The evaluation is performed on an English LVCSR task. Trained on the raw time signal, the convolutional layers allow to reduce the WER on the test set from 25.5% to 23.4%, compared to an MFCC based result of 22.1% using fully connected layers. Index Terms: acoustic modeling, raw time signal, convolutional neural networks Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2015 | Multilingual features based keyword search for very low-resource languagesabstractIn this paper we describe RWTH Aachen’s system for keyword search (KWS) with very limited amount of transcribed audio data available in the target language. This setting has become this year’s primary condition within the Babel project [1], seeking to minimize the amount of human effort while retaining a reasonable KWS performance. Thus the highlights presented in this paper include graphemic acoustic modeling; multilingual features trained on language data from the previous project periods; comparison of tandem and hybrid DNN-HMM acoustic models; processing of large amounts of text data available on the web and the morphological KWS based on automatically derived word fragments. The evaluation is performed using two training sets for each of the six current project period’s languages ‐ full language pack (FLP), consisting of 30 hours and very limited language pack (VLLP), comprising less than 3 hours of transcribed audio data. We put our focus on the latter of the two, which is clearly more challenging. The methods described in this work allowed us to exceed 0.3 MTWV on five out of six languages using development queries. Index Terms: acoustic modeling, keyword search, graphemic, multilingual, neural networks, semi-supervised learning Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2015 | Improvements in RWTH LVCSR evaluation systems for Polish, Portuguese, English, urdu, and ArabicabstractIn this work, Portuguese, Polish, English, Urdu, and Arabic automatic speech recognition evaluation systems developed by the RWTH Aachen University are presented. Our LVCSR systems focus on various domains like broadcast news, spontaneous speech, and podcasts. All these systems but Urdu are used for Euronews and Skynews evaluations as part of the EUBridge project. Our previously developed LVCSR systems were improved using different techniques for the aforementioned languages. Significant improvements are obtained using multilingual tandem and hybrid approaches, minimum phone error training, lexical adaptation, open vocabulary long short term memory language models, maximum entropy language models and confusion-network based system combination. Index Terms: LVCSR, LSTM, open-vocabulary, EU-Bridge M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2014 | Multilingual MRASTA features for low-resource keyword search and speech recognition systemsabstractThis paper investigates the application of hierarchical MRASTA bottleneck (BN) features for under-resourced languages within the IARPA Babel project. Through multilingual training of Multilayer Perceptron (MLP) BN features on five languages (Cantonese, Pashto, Tagalog, Turkish, and Vietnamese), we could end up in a single feature stream which is more beneficial to all languages than the unilingual features. In the case of balanced corpus sizes, the multilingual BN features improve the automatic speech recognition (ASR) performance by 3-5% and the keyword search (KWS) by 3-10% relative for both limited (LLP) and full language packs (FLP). Borrowing orders of magnitude more data from non-target FLPs, the recognition error rate is reduced by 8-10%, and the spoken term detection is improved by over 40% relative on Vietnamese and Pashto LLP. Aiming at the fast development of acoustic models, cross-lingual transfer of multilingually ”pretrained” BN features for a new language is also investigated. Without the need of any MLP training on the new language, the ported BN features performed similarly to the unilingual features on FLP and significantly better on LLP. Results also show that a simple fine-tuning step on the new language is enough to achieve comparable KWS and ASR performance to that system where the target language is also involved in the time-consuming multilingual training. Zoltán Tüske, David Nolden, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2014 | The RWTH English lecture recognition systemabstractIn this paper, we describe the RWTH speech recognition system for English lectures developed within the Translectures project. A difficulty in the development of an English lectures recognition system, is the high ratio of non-native speakers. We address this problem by using very effective deep bottleneck features trained on multilingual data. The acoustic model is trained on large amounts of data from different domains and with different dialects. Large improvements are obtained from unsupervised acoustic adaptation. Another challenge is the frequent use of technical terms and the wide range of topics. In our recognition system, slides, which are attached to most lectures, are used for improving lexical coverage and language model adaptation. Simon Wiesler, Kazuki Irie, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 3 |
| 2014 | RWTH LVCSR systems for quaero and EU-bridge: German, Polish, Spanish and PortugueseabstractIn this paper, German, Polish, Spanish, and Portuguese large vocabulary continuous speech recognition (LVCSR) systems developed by the RWTH Aachen University are presented.All the above mentioned systems for the aforementioned languages are used for the Quaero and EU-Bridge project evaluations.The LVCSR systems developed for these competitive evaluations focus on various domains like broadcast news, podcasts and lecture domain.Transcription of the speech for these tasks is challenging due to huge variability in the acoustic conditions and a significant portion of audio data includes spontaneous speech.Good improvements are obtained using stateof-the-art multilingual bottleneck features, minimum phone error trained acoustic models, language model (LM) adaptation and confusion-network based system combination.In addition, an open vocabulary approach using morphemic units is investigated along with the LM adaptation for the German LVCSR. M. Ali Basha Shaik, Zoltán Tüske, Muhammad Ali Tahir, Markus Nußbaum-Thom, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2014 | Lattice decoding and rescoring with long-Span neural network language modelsabstractWith long-span neural network language models, considerable improvements have been obtained in speech recognition. However, it is difficult to apply these models if the underlying search space is large. In this paper, we combine previous work on lattice decoding with long short-term memory (LSTM) neural network language models. By adding refined pruning techniques, we are able to reduce the search effort by a factor of three. Furthermore, we introduce two novel approximations for full lattice rescoring, which opens the potential of lattice-based speech recognition techniques. Compared to 1000-best lists, we find that we can increase the word error rate improvements obtained with LSTMs from 8.2 % to 10.7 % relative over a stateof-the-art baseline, while the resulting lattices are even considerably smaller. In addition, we investigate the use of LSTMs for Babel Assamese keyword search, obtaining significant improvements of 2.5 % relative. Martin Sundermeyer, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2014 | Data augmentation, feature combination, and multilingual neural networks to improve ASR and KWS performance for low-resource languagesabstractThis paper presents the progress of acoustic models for lowresourced languages (Assamese, Bengali, Haitian Creole, Lao, Zulu) developed within the second evaluation campaign of the IARPA Babel project.This year, the main focus of the project is put on training high-performing automatic speech recognition (ASR) and keyword search (KWS) systems from language resources limited to about 10 hours of transcribed speech data.Optimizing the structure of Multilayer Perceptron (MLP) based feature extraction and switching from the sigmoid activation function to rectified linear units results in about 5% relative improvement over baseline MLP features.Further improvements are obtained when the MLPs are trained on multiple feature streams and by exploiting label preserving data augmentation techniques like vocal tract length perturbation.Systematic application of these methods allows to improve the unilingual systems by 4-6% absolute in WER and 0.064-0.105absolute in MTWV.Transfer and adaptation of multilingually trained MLPs lead to additional gains, clearly exceeding the project goal of 0.3 MTWV even when only the limited language pack of the target language is used. Zoltán Tüske, Pavel Golik, David Nolden, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2014 | Acoustic modeling with deep neural networks using raw time signal for LVCSR
Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2013 | Investigation on cross- and multilingual MLP features under matched and mismatched acoustical conditionsabstractIn this paper, Multi Layer Perceptron (MLP) based multilingual bottleneck features are investigated for acoustic modeling in three languages - German, French, and US English. We use a modified training algorithm to handle the multilingual training scenario without having to explicitly map the phonemes to a common phoneme set. Furthermore, the cross-lingual portability of bottleneck features between the three languages are also investigated. Single pass recognition experiments on large vocabulary SMS dictation task indicate that (1) multilingual bottleneck features yield significantly lower word error rates compared to standard MFCC features (2) multilingual bottleneck features are superior to monolingual bottleneck features trained for the target language with limited training data, and (3) multilingual bottleneck features are beneficial in training acoustic models in a low resource language where only mismatched training data is available-by exploiting the more matched training data from other languages. Zoltán Tüske, Joel Pinto, Daniel Willett, Ralf Schlüter |
ICASSP | 1 |
| 2013 | Deep hierarchical bottleneck MRASTA features for LVCSRabstractHierarchical Multi Layer Perceptron (MLP) based long-term feature extraction is optimized for TANDEM connectionist large vocabulary continuous speech recognition (LVCSR) system within the QUAERO project. Training the bottleneck MLP on multi-resolutional RASTA filtered critical band energies, more than 20% relative word error rate (WER) reduction over standard MFCC system is observed after optimizing the number of target labels. Furthermore, introducing a deeper structure in the hierarchical bottleneck processing the relative gain increases to 25%. The final system based on deep bottleneck TANDEM features clearly outperforms the hybrid approach, even if the long-term features are also presented to the deep MLP acoustic model. The results are also verified on evaluation data of the year 2012, and about 20% relative WER improvement over classical cepstral system is measured even after speaker adaptive training. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2013 | Development of the RWTH transcription system for slovenianabstractIn this paper we describe the RWTH automatic speech recognition system for Slovenian developed within the transLectures project.The project aims at supporting the transcription and translation of video lectures freely available on the web.Difficulties arise on all levels of modeling: Slovenian is a morphologically rich language with a high level of inflection (pronunciation model), and a large variety of dialects and recording conditions brings uncertainty into the audio signal (acoustic model).Moreover, the video lectures cover a wide spectrum of topics with a high share of spontaneous speech and technical terms (language model).These issues require application of robust and adaptive methods.Besides the system description, this study mainly focuses on robust acoustic modeling.Building acoustic models from various resources, we also compare the influence of speaker adaptation to different neural network based acoustic features.Systematic application of these methods allows us to reduce the word error rate on the evaluation corpus from 59.2% to 43.4%.We also give a motivation for Slovenian open vocabulary recognition and perform some first steps. Pavel Golik, Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2013 | Multilingual hierarchical MRASTA features for ASRabstractRecently, a multilingual Multi Layer Perceptron (MLP) training method was introduced without having to explicitly map the phonetic units of multiple languages to a common set.This paper further investigates this method using bottleneck (BN) tandem connectionist acoustic modeling for four high-resourced languages -English, French, German, and Polish.Aiming at the improvement of already existing high performing automatic speech recognition (ASR) systems, the multilingual training of the BN-MLP is extended from short-term to hierarchical longterm (multi-resolutional RASTA) feature extraction.Furthermore, deeper structures and context-dependent target labels are also examined.We experimentally demonstrate that a single state-of-the-art BN feature set can be trained for multiple languages, which is superior to the monolingual feature set, and results in significant gains in all the four languages.Studying the scalability of the multilingual BN features, a similar gain is observed in small (50 hours) and in larger scale (300 hours) ASR experiments regardless of the distribution of the data amount between the languages.Using deeper structures, context-dependent targets, and speaker adaptation, the multilingual BN reduces the word error rates by 3-7% relative over the target language BN features and 25-30% over the conventional MFCC system. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2012 | Comparison and combination of different CRBE based MLP features for LVCSRabstractMulti Layer Perceptron (MLP) features extracted from different types of critical band energies (CRBE) - derived from MFCC, GT, and PLP pipeline - are compared on French broadcast news and conversational speech recognition task. Though the MLP structure is kept fixed, ROVER combination of different CRBE based systems leads to 4% relative improvement. Furthermore, aiming at the combination of state-of-the-art features based on various signal analysis methods into one single stream, posterior feature space based combination technique is proposed. The speaker normalized features originated from different CRBEs are merged after additional MLP training by Dempster-Shafer rule. The performance of these posterior features unifying the different CRBE based features is superior to the best single CRBE based posterior features by 6% relative. Further results reveal that the concatenated cepstral and unified posterior features perform nearly as well as the ROVER combination of the different CRBE based systems. Zoltán Tüske, Ralf Schlüter, Hermann Ney |
ICASSP | 1 |
| 2012 | Posterior-Scaled MPE: Novel Discriminative Training CriteriaabstractWe recently discovered novel discriminative training criteria following a principled approach. In this approach training criteria are developed from error bounds on the global error for pattern classification tasks that depend on non-trivial loss functions. Automatic speech recognition (ASR) is a prominent example for such a task depending on the non-trivial Levenshtein loss. In this context, the posterior-scaled Minimum Phoneme Error (MPE) training criterion, which is the state-of-the-art discriminative training criterion in ASR, was shown to be an approximation to one of the novel criteria. Here, we describe the implementation of the posterior-scaled MPE criterion in a transducer-based framework, and compare this criterion to other discriminative training criteria on an ASR task. This comparison indicates that the posterior-scaled MPE criterion performs better than other discriminative criteria including MPE. Index Terms: error bounds, discriminative training criteria, margin, MPE Markus Nußbaum-Thom, Zoltán Tüske, Georg Heigold, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2012 | Context-Dependent MLPs for LVCSR: TANDEM, Hybrid or Both?abstractGaussian Mixture Model (GMM) and Multi Layer Perceptron (MLP) based acoustic models are compared on a French large vocabulary continuous speech recognition (LVCSR) task.In addition to optimizing the output layer size of the MLP, the effect of the deep neural network structure is also investigated.Moreover, using different linear transformations (time derivatives, LDA, CMLLR) on conventional MFCC, the study is also extended to MLP based probabilistic and bottle-neck TANDEM features.Results show that using either the hybrid or bottleneck TANDEM approach leads to similar recognition performance.However, the best performance is achieved when deep MLP acoustic models are trained on concatenated cepstral and context-dependent bottle-neck features.Further experiments reveal the importance of the neighbouring frames in case of MLP based modeling, and that its gain over GMM acoustic models is strongly reduced by more complex features. Zoltán Tüske, Ralf Schlüter, Hermann Ney, Martin Sundermeyer |
INTERSPEECH | 1 |
| 2011 | Non-stationary feature extraction for automatic speech recognitionabstractIn current speech recognition systems mainly Short-Time Fourier Transform based features like MFCC are applied. Dropping the short-time stationarity assumption of the voiced speech, this paper introduces the non-stationary signal analysis into the ASR framework. We present new acoustic features extracted by a pitch-adaptive Gammatone filter bank. The noise robustness was proved on AURORA 2 and 4 tasks, where the proposed features outperform the standard MFCC. Furthermore, successful combination experiments via ROVER indicate the differences between the new features and MFCC. Zoltán Tüske, Pavel Golik, Ralf Schlüter, Friedhelm R. Drepper |
ICASSP | 1 |
| 2011 | A Study on Speaker Normalized MLP Features in LVCSRabstractDifferent normalization methods are applied in recent Large Vocabulary Continuous Speech Recognition Systems (LVCSR) to reduce the influence of speaker variability on the acoustic models.In this paper we investigate the use of Vocal Tract Length Normalization (VTLN) and Speaker Adaptive Training (SAT) in Multi Layer Perceptron (MLP) feature extraction on an English task.We achieve significant improvements by each normalization method and we gain further by stacking the normalizations.Studying features transformed by Constrained Maximum Likelihood Linear Regression (CMLLR) based SAT as possible input for MLP, further experiments show that MLP could not consistently take advantage of SAT as it does in case of VTLN. Zoltán Tüske, Christian Plahl, Ralf Schlüter |
INTERSPEECH | 1 |
| 2010 | Improved Recognition of Spontaneous Hungarian Speech - Morphological and Acoustic Modeling Techniques for a Less Resourced TaskabstractVarious morphological and acoustic modeling techniques are evaluated on a less resourced, spontaneous Hungarian large-vocabulary continuous speech recognition (LVCSR) task. Among morphologically rich languages, Hungarian is known for its agglutinative, inflective nature that increases the data sparseness caused by a relatively small training database. Although Hungarian spelling is considered as simple phonological, a large part of the corpus is covered by words pronounced in multiple, phonemically different ways. Data-driven and language specific knowledge supported vocabulary decomposition methods are investigated in combination with phoneme- and grapheme-based acoustic modeling techniques on the given task. Word baseline and morph-based advanced baseline results are significantly outperformed by using both statistical and grammatical vocabulary decomposition methods. Although the discussed morph-based techniques recognize a significant amount of out of vocabulary words, the improvements are due not to this fact but to the reduction of insertion errors. Applying grapheme-based acoustic models instead of phoneme-based models causes no severe recognition performance deteriorations. Moreover, a fully data-driven acoustic modeling technique along with a statistical morphological modeling approach provides the best performance on the most difficult test set. The overall best speech recognition performance is obtained by using a novel word to morph decomposition technique that combines grammatical and unsupervised statistical segmentation algorithms. The improvement achieved by the proposed technique is stable across acoustic modeling approaches and larger with speaker adaptation. Péter Mihajlik, Zoltán Tüske, Balázs Tarján, Bottyán Németh, Tibor Fegyó |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Investigation of morph-based speech recognition improvements across speech genres
Péter Mihajlik, Balázs Tarján, Zoltán Tüske, Tibor Fegyó |
INTERSPEECH | 3 |
| 2007 | A morpho-graphemic approach for the recognition of spontaneous speech in agglutinative languages - like HungarianabstractA coupled acoustic- and language-modeling approach is presented for the recognition of spontaneous speech primarily in agglutinative languages. The effectiveness of the approach in large vocabulary spontaneous speech recognition is demonstrated on the Hungarian MALACH corpus. The derivation of morphs from word forms is based on a statistical morphological segmentation tool while the mapping of morphs into graphemes is obtained trivially by splitting each morph into individual letters. Using morphs instead of words in language modeling gives significant WER reductions in case of both phoneme- and grapheme-based acoustic modeling. The improvements are larger after speaker adaptation of the acoustic models. In conclusion, morphophonemic and the proposed morpho-graphemic ASR approaches yield the same best WERs, which are significantly lower than the word-based baselines but essentially without language dependent rules or pronunciation dictionaries in the latter case. Index Terms: spontaneous speech recognition, morphology. Péter Mihajlik, Tibor Fegyó, Zoltán Tüske, Pavel Ircing |
INTERSPEECH | 3 |
| 2005 | Evaluation and optimization of noise robust front-end technologies for the automatic recognition of Hungarian telephone speechabstractIn this paper a variety of front-end configurations are evaluated on Hungarian telephone speech databases. Our aim was to measure directly the efficiency of the front-ends on real noisy and normal speech data. As a baseline the ETSI ADSR standard front-end is used. Some simplification on the standard is introduced resulting in better performance on our databases than the original front-end in terms of both speed and recognition rate. Besides, another recently proposed feature extraction approach is also investigated. Finally the effect of the novel voice activity detection approach is evaluated. The best front-end configuration augmented with this voice activity detector outperformed significantly the baseline in each recognition test and by 24,7% relative in average. Péter Mihajlik, Zoltán Tobler, Zoltán Tüske, Géza Gordos |
INTERSPEECH | 3 |
| 2005 | Robust voice activity detection based on the entropy of noise-suppressed spectrumabstractA novel noise robust voice activity detection approach is introduced. The novelty of the method that it uses noise suppressed spectrum of the input signal for spectral entropy calculation. As a result excellent end-pointing performance is observed based on predefined global entropy threshold and time constraints. The effect of frame dropping controlled by the proposed algorithm was investigated on the accuracy of automatic speech recognition. The experiments were performed on Hungarian publicly available noisy and normal telephony speech databases. The relative improvement due to dropping of non-speech frames was positive in all test configurations with a maximum of 29,5%. Besides, in average more than 50% of the frames were dropped. Zoltán Tüske, Péter Mihajlik, Zoltán Tobler, Tibor Fegyó |
INTERSPEECH | 1 |