EDBT 2026 Demo / reviewers in the wild / expert
Björn Hoffmeister
dblp:32/1010
· DBLP profile ↗
39ranked-venue papers
7as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 27 · 6 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 67% Question answering and dialogue systems · 26% Machine translation · 5% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.9 | 2 | 2024 | An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems · ICML 2024 WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
contextual ASR |
0.8 | 1 | 2024 | An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems · ICML 2024 |
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems |
0.8 | 1 | 2024 | An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems · ICML 2024 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Machine translation
system combination |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis › search and decoding
weighted finite-state transducer decoding |
0.1 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
0.0 | 1 | 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding · IEEE Trans. Speech Audio Process. 2012 |
Methods — techniques the papers use, named apart from their topics
student-teacher learning · 0.8hard negative mining · 0.8contrastive self-supervision · 0.8weighted finite-state transducer · 0.1levenshtein distance · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech RecognitionabstractWhile word error rates of automatic speech recognition (ASR) systems have consistently fallen, natural language understanding (NLU) applications built on top of ASR systems still attribute significant numbers of failures to low-quality speech recognition results. Existing assistant systems collect large numbers of these unsuccessful interactions, but these systems usually fail to learn from these interactions, even in an offline fashion. In this work, we introduce CLC: Contrastive Learning for Conversations, a family of methods for contrastive fine-tuning of models in a self-supervised fashion, making use of easily detectable artifacts in unsuccessful conversations with assistants. We demonstrate that our CLC family of approaches can improve the performance of ASR models on OD3, a new public large-scale semi-synthetic meta-dataset of audio task-oriented dialogues, by up to 19.2%. These gains transfer to real-world systems as well, where we show that CLC can help to improve performance by up to 6.7% over baselines.1 David M. Chan, Shalini Ghosh, Hitesh Tulsiani, Ariya Rastrow, Björn Hoffmeister |
ICASSP | 5 |
| 2024 | An Efficient Self-Learning Framework For Interactive Spoken Dialog SystemsabstractDialog systems, such as voice assistants, are expected to engage with users in complex, evolving conversations. Unfortunately, traditional automatic speech recognition (ASR) systems deployed in such applications are usually trained to recognize each turn independently and lack the ability to adapt to the conversational context or incorporate user feedback. In this work, we introduce a general framework for ASR in dialog systems that can go beyond learning from single-turn utterances and learn over time how to adapt to both explicit supervision and implicit user feedback present in multi-turn conversations. We accomplish that by leveraging advances in student-teacher learning and context-aware dialog processing, and designing contrastive self-supervision approaches with Ohm, a new online hard-negative mining approach. We show that leveraging our new framework compared to traditional training leads to relative WER reductions of close to 10% in real-world dialog systems, and up to 26% on public synthetic data. Hitesh Tulsiani, David M. Chan, Shalini Ghosh, Garima Lalwani, Prabhat Pandey, Ankish Bansal, Sri Garimella, Ariya Rastrow, Björn Hoffmeister |
ICML | 9 |
| 2023 | Domain Adaptation with External Off-Policy Acoustic Catalogs for Scalable Contextual End-to-End Automated Speech RecognitionabstractDespite improvements to the generalization performance of automated speech recognition (ASR) models, specializing ASR models for downstream tasks remains a challenging task, primarily due to reduced data availability (necessitating increased data collection), and rapidly shifting data distributions (requiring more frequent model fine-tuning). In this work, we investigate the potential of leveraging external knowledge, particularly through off-policy generated text-to-speech key-value stores, to allow for flexible post-training adaptation to new data distributions. In our approach, audio embeddings captured from text-to-speech are used, along with semantic text embeddings, to bias ASR via an approximate k-nearest-neighbor (KNN) based attentive fusion step. Our experiments on LibiriSpeech and in-house voice assistant/search datasets show that the proposed approach can reduce domain adaptation time by up to 1K GPU-hours while providing up to 3% WER improvement compared to a fine-tuning baseline, suggesting a promising approach for adapting production ASR systems in challenging zero and few-shot scenarios. David M. Chan, Shalini Ghosh, Ariya Rastrow, Björn Hoffmeister |
ICASSP | 4 |
| 2022 | Multi-Modal Pre-Training for Automated Speech RecognitionabstractTraditionally, research in automated speech recognition has focused on local-first encoding of audio representations to predict the spoken phonemes in an utterance. Unfortunately, approaches relying on such hyper-local information tend to be vulnerable to both local-level corruption (such as audio-frame drops, or loud noises) and global-level noise (such as environmental noise, or background noise) that has not been seen during training. In this work, we introduce a novel approach that leverages a self-supervised learning technique based on masked language modeling to compute a global, multi-modal encoding of the environment in which the utterance occurs. We then use a new deep-fusion framework to integrate this global context into a traditional ASR method, and demonstrate that the resulting method can outperform baseline methods by up to 7% on Librispeech; gains on internal datasets range from 6% (on larger models) to 45% (on smaller models). David M. Chan, Shalini Ghosh, Debmalya Chakrabarty, Björn Hoffmeister |
ICASSP | 4 |
| 2020 | DiPCo - Dinner Party CorpusabstractWe present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment. The corpus was created by recording multiple groups of four Amazon employee volunteers having a natural conversation in English around a dining table. The participants were recorded by a single-channel close-talk microphone and by five far-field 7-microphone array devices positioned at different locations in the recording room. The dataset contains the audio recordings and human labeled transcripts of a total of 10 sessions with a duration between 15 and 45 minutes. The corpus was created to advance in the field of noise robust and distant speech processing and is intended to serve as a public research and benchmarking data set. Maarten Van Segbroeck, Ahmed Zaid, Ksenia Kutsenko, Cirenia Huerta, Tinh Nguyen, Xuewen Luo, Björn Hoffmeister, Jan Trmal, Maurizio Omologo, Roland Maas |
INTERSPEECH | 7 |
| 2019 | Multi-geometry Spatial Acoustic Modeling for Distant Speech RecognitionabstractThe use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array geometry mismatch between design and test conditions. Moreover, such speech enhancement techniques do not always yield ASR accuracy improvement due to the difference between speech enhancement and ASR optimization objectives. In this work, we propose to unify an acoustic model framework by optimizing spatial filtering and long short-term memory (LSTM) layers from multi-channel (MC) input. Our acoustic model subsumes beamformers with multiple types of array geometry. In contrast to deep clustering methods that treat a neural network as a black box tool, the network encoding the spatial filters can process streaming audio data in real time without the accumulation of target signal statistics. We demonstrate the effectiveness of such MC neural networks through ASR experiments on the real-world far-field data. We show that our two-channel acoustic model can on average reduce word error rates (WERs) by 13.4 and 12.7% compared to a single channel ASR system with the log-mel filter bank energy (LFBE) feature under the matched and mismatched microphone placement conditions, respectively. Our result also shows that our two-channel network achieves a relative WER reduction of over 7.0% compared to conventional beamforming with seven microphones overall. Ken'ichi Kumatani, Minhua Wu, Shiva Sundaram, Nikko Strom, Björn Hoffmeister |
ICASSP | 5 |
| 2019 | Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student LearningabstractFor real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we apply a logits selection method which only preserves the k highest values to prevent wrong emphasis of knowledge from the teacher and to reduce bandwidth needed for transferring data. We incorporate up to 8000 hours of untranscribed data for training and present our results on sequence trained models apart from cross entropy trained ones. The best sequence trained student model yields relative word error rate (WER) reductions of approximately 10.1%, 28.7% and 19.6% on our clean, simulated noisy and real test sets respectively comparing to a sequence trained teacher. Ladislav Mosner, Minhua Wu, Anirudh Raju, Sree Hari Krishnan Parthasarathi, Ken'ichi Kumatani, Shiva Sundaram, Roland Maas, Björn Hoffmeister |
ICASSP | 8 |
| 2019 | End-to-end Anchored Speech RecognitionabstractVoice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to be ignored. We propose two end-to-end models to tackle this problem with information extracted from the anchored segment. The anchored segment refers to the wake-up word part of an audio stream, which contains valuable speaker information that can be used to suppress interfering speech and background noise. The first method is called Multi-source Attention where the attention mechanism takes both the speaker information and decoder state into consideration. The second method directly learns a frame-level mask on top of the encoder output. We also explore a multi-task learning setup where we use the ground truth of the mask to guide the learner. Given that audio data with interfering speech is rare in our training data, we also propose a way to synthesize "noisy" speech from "clean" speech to mitigate the mismatch between training and test data. Our proposed methods show up to 15% relative reduction in WER for Amazon Alexa live data with interfering background speech without significantly degrading on clean speech. Yiming Wang 0006, I-Fan Chen, Yuzong Liu, Tongfei Chen, Björn Hoffmeister |
ICASSP | 6 |
| 2019 | Frequency Domain Multi-channel Acoustic Modeling for Distant Speech RecognitionabstractConventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques do not always yield ASR accuracy improvement because the optimization criterion for speech enhancement is not directly relevant to the ASR objective. In this work, we develop new acoustic modeling techniques that optimize spatial filtering and long short-term memory (LSTM) layers from multi-channel (MC) input based on an ASR criterion directly. In contrast to conventional methods, we incorporate array processing knowledge into the acoustic model. Moreover, we initialize the network with beamformers’ coefficients. We investigate effects of such MC neural networks through ASR experiments on the real-world far-field data where users are interacting with an ASR system in uncontrolled acoustic environments. We show that our MC acoustic model can reduce a word error rate (WER) by 16.5% compared to a single channel ASR system with the traditional log-mel filter bank energy (LFBE) feature on average. Our result also shows that our network with the spatial filtering layer on two-channel input achieves a relative WER reduction of 9.5% compared to conventional beamforming with seven microphones. Minhua Wu, Ken'ichi Kumatani, Shiva Sundaram, Nikko Strom, Björn Hoffmeister |
ICASSP | 5 |
| 2019 | A Study for Improving Device-Directed Speech Detection Toward Frictionless Human-Machine Interaction
Che-Wei Huang, Roland Maas, Sri Harish Reddy Mallidi, Björn Hoffmeister |
INTERSPEECH | 4 |
| 2019 | Improving ASR Confidence Scores for Alexa Using Acoustic and Hypothesis Embeddings
Prakhar Swarup, Roland Maas, Sri Garimella, Sri Harish Reddy Mallidi, Björn Hoffmeister |
INTERSPEECH | 5 |
| 2018 | Combining Acoustic Embeddings and Decoding Features for End-of-Utterance Detection in Real-Time Far-Field Speech Recognition SystemsabstractWe present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on l-best recognition hypotheses of the automatic speech recognition (ASR) decoder, and a feedforward deep neural network (DNN) combining embeddings derived from both LSTMs with pause duration features from the ASR decoder. At inference time, lower and upper latency (pause duration) bounds act as safeguards. Within the latency bounds, the utterance end-point is triggered as soon as the DNN posterior reaches a tuned threshold. Our experimental evaluation is carried out on real recordings of natural human interactions with voice-controlled far-field devices. We show that the acoustic embeddings are the single most powerful feature and particularly suitable for cross-lingual applications. We furthermore show the benefit of ASR decoder features, especially as a low cost alternative to ASR hypothesis em-beddings. Roland Maas, Ariya Rastrow, Chengyuan Ma, Guitang Lan, Kyle Goehner, Gautam Tiwari, Shaun Joseph, Björn Hoffmeister |
ICASSP | 8 |
| 2018 | Monophone-Based Background Modeling for Two-Stage On-Device Wake Word DetectionabstractAccurate on-device wake word detection is crucial to products with far-field voice control such as the Amazon Echo. It is quite challenging to build a wake word system with both low False Reject Rate (FRR) and low False Alarm Rate (FAR) in real scenarios where there are various types of background speech, music or noise, especially when computational resources on the device is limited. In this paper, we introduce a two-stage wake word system based on Deep Neural Network (DNN) acoustic modeling, propose a new way to model the non-keyword background events using monophone-based units and present how richer information can be extracted from those monophone units for final wake word detection. Under the new system, we could get around 16% relative reduction in FRR when fixing the false alarm level, and about 37% relative reduction in FAR on the other hand if we maintain the miss rate. For the 2nd stage classifier itself, it is able to reduce the false alarm rate relatively by about 67% on top of 1st stage hypothesis with very few computational resources. Minhua Wu, Sankaran Panchapagesan, Ming Sun 0007, Jiacheng Gu, Ryan Thomas, Shiv Vitaladevuni, Björn Hoffmeister, Arindam Mandal |
ICASSP | 7 |
| 2018 | Device-directed Utterance DetectionabstractIn this work, we propose a classifier for distinguishing device-directed queries from background speech in the context of interactions with voice assistants.Applications include rejection of false wake-ups or unintended interactions as well as enabling wake-word free followup queries.Consider the example interaction: "Computer, play music", "Computer, reduce the volume".In this interaction, the user needs to repeat the wake-word (Computer) for the second query.To allow for more natural interactions, the device could immediately re-enter listening state after the first query (without wake-word repetition) and accept or reject a potential follow-up as device-directed or background speech.The proposed model consists of two long short-term memory (LSTM) neural networks trained on acoustic features and automatic speech recognition (ASR) 1-best hypotheses, respectively.A feed-forward deep neural network (DNN) is then trained to combine the acoustic and 1-best embeddings, derived from the LSTMs, with features from the ASR decoder.Experimental results show that ASR decoder, acoustic embeddings, and 1-best embeddings yield an equal-error-rate (EER) of 9.3 %, 10.9 % and 20.1 %, respectively.Combination of the features resulted in a 44 % relative improvement and a final EER of 5.2 %. Sri Harish Reddy Mallidi, Roland Maas, Kyle Goehner, Ariya Rastrow, Spyridon Matsoukas, Björn Hoffmeister |
INTERSPEECH | 6 |
| 2018 | Scalable Language Model Adaptation for Spoken Dialogue SystemsabstractLanguage models (LM) for interactive speech recognition systems are trained on large amounts of data and the model parameters are optimized on past user data. New application intents and interaction types are released for these systems over time, imposing challenges to adapt the LMs since the existing training data is no longer sufficient to model the future user interactions. It is unclear how to adapt LMs to new application intents without degrading the performance on existing applications. In this paper, we propose a solution to (a) estimate n-gram counts directly from the hand-written grammar for training LMs and (b) use constrained optimization to optimize the system parameters for future use cases, while not degrading the performance on past usage. We evaluated our approach on new applications intents for a personal assistant system and find that the adaptation improves the word error rate by up to 15% on new applications even when there is no adaptation data available for an application. Ankur Gandhe, Ariya Rastrow, Björn Hoffmeister |
SLT | 3 |
| 2018 | LSTM-Based Whisper DetectionabstractThis article presents a whisper speech detector in the far-field domain. The proposed system consists of a long-short term memory (LSTM) neural network trained on log-filterbank energy (LFBE) acoustic features. This model is trained and evaluated on recordings of human interactions with voice-controlled, far-field devices in whisper and normal phonation modes. We compare multiple inference approaches for utterance-level classification by examining trajectories of the LSTM posteriors. In addition, we engineer a set of features based on the signal characteristics inherent to whisper speech, and evaluate their effectiveness in further separating whisper from normal speech. A benchmarking of these features using multilayer perceptrons (MLP) and LSTMs suggests that the proposed features, in combination with LFBE features, can help us further improve our classifiers. We prove that, with enough data, the LSTM model is indeed as capable of learning whisper characteristics from LFBE features alone compared to a simpler MLP model that uses both LFBE and features engineered for separating whisper and normal speech. In addition, we prove that the LSTM classifiers accuracy can be further improved with the incorporation of the proposed engineered features. Zeynab Raeesy, Kellen Gillespie, Chengyuan Ma, Thomas Drugman, Jiacheng Gu, Roland Maas, Ariya Rastrow, Björn Hoffmeister |
SLT | 8 |
| 2017 | Robust Speech Recognition via Anchor Word Representations
Brian John King, I-Fan Chen, Yonatan Vaizman, Yuzong Liu, Roland Maas, Sree Hari Krishnan Parthasarathi, Björn Hoffmeister |
INTERSPEECH | 7 |
| 2017 | Zero-Shot Learning Across Heterogeneous Overlapping Domains
Anjishnu Kumar, Pavankumar Reddy Muddireddy, Markus Dreyer, Björn Hoffmeister |
INTERSPEECH | 4 |
| 2017 | Domain-Specific Utterance End-Point Detection for Speech Recognition
Roland Maas, Ariya Rastrow, Kyle Goehner, Gautam Tiwari, Shaun Joseph, Björn Hoffmeister |
INTERSPEECH | 6 |
| 2016 | LatticeRnn: Recurrent Neural Networks Over Lattices
Faisal Ladhak, Ankur Gandhe, Markus Dreyer, Lambert Mathias, Ariya Rastrow, Björn Hoffmeister |
INTERSPEECH | 6 |
| 2016 | Anchored Speech Detection
Roland Maas, Sree Hari Krishnan Parthasarathi, Brian John King, Ruitong Huang, Björn Hoffmeister |
INTERSPEECH | 5 |
| 2016 | Multi-Task Learning and Weighted Cross-Entropy for DNN-Based Keyword Spotting
Sankaran Panchapagesan, Ming Sun 0007, Aparna Khare, Spyridon Matsoukas, Arindam Mandal, Björn Hoffmeister, Shiv Vitaladevuni |
INTERSPEECH | 6 |
| 2015 | Model Shrinking for Embedded Keyword SpottingabstractIn this paper we present two approaches to improve computational efficiency of a keyword spotting system running on a resource constrained device. This embedded keyword spotting system detects a pre-specified keyword in real time at low cost of CPU and memory. Our system is a two stage cascade. The first stage extracts keyword hypotheses from input audio streams. After the first stage is triggered, hand-crafted features are extracted from the keyword hypothesis and fed to a support vector machine (SVM) classifier on the second stage. This paper focuses on improving the computational efficiency of the second stage SVM classifier. More specifically, select a subset of feature dimensions and merge the SVM classifier to a smaller size, while maintaining the keyword spotting performance. Experimental results indicate that we can remove more than 36% of the non-discriminative SVM features, and reduce the number of support vectors by more than 60% without significant performance degradation. This results in more than 15% relative reduction in CPU utilization. Ming Sun 0007, Varun K. Nagaraja, Björn Hoffmeister, Shiv Vitaladevuni |
ICMLA | 3 |
| 2015 | Robust i-vector based adaptation of DNN acoustic model for speech recognitionabstractIn the past, conventional i-vectors based on a Universal Background Model (UBM) have been successfully used as input features to adapt a Deep Neural Network (DNN) Acoustic Model (AM) for Automatic Speech Recognition (ASR). In contrast, this paper introduces Hidden Markov Model (HMM) based ivectors that use HMM state alignment information from an ASR system for estimating i-vectors. Further, we propose passing these HMM based i-vectors though an explicit non-linear hidden layer of a DNN before combining them with standard acoustic features, such as log filter bank energies (LFBEs). To improve robustness to mismatched adaptation data, we also propose estimating i-vectors in a causal fashion for training the DNN, restricting the connectivity among hidden nodes in the DNN and applying a max-pool non-linearity at selected hidden nodes. In our experiments, these techniques yield about 5-7% relative word error rate (WER) improvement over the baseline speaker independent system in matched condition, and a substantial WER reduction for mismatched adaptation data. Sri Garimella, Arindam Mandal, Nikko Strom, Björn Hoffmeister, Spyridon Matsoukas, Sree Hari Krishnan Parthasarathi |
INTERSPEECH | 4 |
| 2015 | Accurate endpointing with expected pause duration
Baiyang Liu, Björn Hoffmeister, Ariya Rastrow |
INTERSPEECH | 2 |
| 2015 | fMLLR based feature-space speaker adaptation of DNN acoustic modelsabstractWe investigate the problem of speaker adaptation of DNN acoustic models in two settings: the traditional unsupervised adaptation and a supervised adaptation (SuA) where a few minutes of transcribed speech is available. SuA presents additional difficulties when a test speaker’s adaptation information does not match the registered speaker’s information. Employing feature-space maximum likelihood linear regression (fMLLR) transformed features as side-information to the DNN, we reintroduce some classical ideas for combining adapted and unadapted features: early and late fusion methods, as well as the estimation of the fMLLR transforms using simple target models (STM). Results show that early fusion helps DNNs generalize better when features are combined after a non-linear bottleneck layer, while late fusion improves robustness, specifically in mismatched cases. STM give consistent improvements in both settings. Sree Hari Krishnan Parthasarathi, Björn Hoffmeister, Spyridon Matsoukas, Arindam Mandal, Nikko Strom, Sri Garimella |
INTERSPEECH | 2 |
| 2012 | WFST Enabled Solutions to ASR Problems: Beyond HMM DecodingabstractDuring the last decade, weighted finite-state transducers (WFSTs) have become popular in speech recognition. While their main field of application remains hidden Markov model (HMM) decoding, the WFST framework is now also seen as a brick in solutions to many other central problems in automatic speech recognition (ASR). These solutions are less known, and this work aims at giving an overview of the applications of WFSTs in large-vocabulary continuous speech recognition (LVCSR) besides HMM decoding: discriminative acoustic model training, Bayes risk decoding, and system combination. The application of the WFST framework has a big practical impact: we show how the framework helps to structure problems, to develop generic solutions, and to delegate complex computations to WFST toolkits. In this paper, we review the literature, discuss existing approaches, and provide new insights into WFST enabled solutions. We also present a novel, purely WFST-based algorithm for computing the exact Bayes risk hypothesis from a lattice with the Levenshtein distance as loss function. We present the problems and their solutions in a unified framework and discuss the advantages and limits of using WFSTs. We do not provide new experimental results, but refer to the existing literature. Our work helps to identify where and how the transducer framework can contribute to a compact and generic solution to LVCSR problems. Björn Hoffmeister, Georg Heigold, David Rybach, Ralf Schlüter, Hermann Ney |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Log-linear model combination with word-dependent scaling factorsabstractLog-linear model combination is the standard approach in LVCSR to combine several knowledge sources, usually an acoustic and a language model.Instead of using a single scaling factor per knowledge source, we make the scaling factor wordand pronunciation-dependent.In this work, we combine three acoustic models, a pronunciation model, and a language model for a Mandarin BN/BC task.The achieved error rate reduction of 2% relative is small but consistent for two test sets.An analysis of the results shows that the major contribution comes from the improved interdependency of language and acoustic model. Björn Hoffmeister, Ruoying Liang, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2009 | Bayes risk approximations using time overlap with an application to system combinationabstractThe computation of the Minimum Bayes Risk (MBR) decoding rule for word lattices needs approximations. We investigate a class of approximations where the Levenshtein alignment is approximated under the condition that competing lattice arcs overlap in time. The approximations have their origins in MBR decoding and in discriminative training. We develop modified versions and propose a new, conceptually extremely simple confusion network algorithm. The MBR decoding rule is extended to scope with several lattices, which enables us to apply all the investigated approximations to system combination. All approximations are tested on a Mandarin and on an English LVCSR task for a single system and for system combination. The new methods are competitive in error rate and show some advantages over the standard approaches to MBR decoding. Index Terms: speech recognition, minimum bayes risk, confusion network, system combination, discriminative training Björn Hoffmeister, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2009 | Development of the GALE 2008 Mandarin LVCSR systemabstractThis paper describes the current improvements of the RWTH Mandarin LVCSR system.We introduce vocal tract length normalization for the Gammatone features and present comparable results for Gammatone based feature extraction and classical feature extraction.In order to benefit from the huge amount of data of 1600h available in the GALE project we have trained the acoustic models up to 8M Gaussians.We present detailed character error rates for the different number of Gaussians.Different kinds of systems are developed and a two stage decoding framework is applied, which uses cross-adaptation and a subsequent lattice-based system combination.In addition to various acoustic front-ends, these systems use different kinds of neural network toneme posterior features.We present detailed recognition results of the development cycle and the different acoustic front-ends of the systems.Finally, we compare the ultimate evaluation system to our last years system and can report a 10% relative improvement. Christian Plahl, Björn Hoffmeister, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2009 | The RWTH aachen university open source speech recognition systemabstractWe announce the public availability of the RWTH Aachen University speech recognition toolkit.The toolkit includes state of the art speech recognition technology for acoustic model training and decoding.Speaker adaptation, speaker adaptive training, unsupervised training, a finite state automata library, and an efficient tree search decoder are notable components.Comprehensive documentation, example setups for training and recognition, and a tutorial are provided to support newcomers. David Rybach, Christian Gollan, Georg Heigold, Björn Hoffmeister, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 4 |
| 2008 | iCNC and iROVER: the limits of improving system combination with classification?abstractWe show how ROVER and confusion network combination (CNC) can be improved with classification. The general idea of improving combination with classification is that each word is assigned to a certain location and at each location a classifier decides which of the provided alternatives is most likely correct. We investigate four variations of this idea and three different classifiers, which are trained on various features derived from ASR lattices. For our experiments, we use highly optimized ROVER and CNC systems as baseline, which already give a relative reduction in WER of more than 20 % for the TC-STAR 2007 English task. With our methods we can further improve the result of the corresponding standard combination method. Index Terms: speech recognition, system combination 1. Björn Hoffmeister, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2008 | Spoken language translation systems ************ ASR word lattice translation with exhaustive reordering is possible
Evgeny Matusov, Björn Hoffmeister, Hermann Ney |
INTERSPEECH | 2 |
| 2008 | Recent improvements of the RWTH GALE Mandarin LVCSR systemabstractThis paper describes the current improvements of the RWTH Mandarin LVCSR system. We introduce a new reduced toneme set developed at RWTH. We are using different toneme sets and pronunciation lexica. For the purpose of discriminative training we will show a fast way to transform word lattices between systems using different toneme sets and pronunciation lexica. In addition to various acoustic front-ends, the current systems use different kinds of neural network toneme posterior features. While different kinds of systems are developed, a two stage decoding framework for combining these systems is applied. We show detailed recognition results of the development cycle of the systems. Finally, two methods to integrate tonal features are compared. Christian Plahl, Björn Hoffmeister, Mei-Yuh Hwang, Danju Lu, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 2 |
| 2007 | Development of the 2007 RWTH Mandarin LVCSR systemabstractThis paper describes the development of the RWTH Mandarin LVCSR system. Different acoustic front-ends together with multiple system cross-adaptation are used in a two stage decoding framework. We describe the system in detail and present systematic recognition results. Especially, we compare a variety of approaches for cross-adapting to multiple systems. During the development we did a comparative study on different methods for integrating tone and phoneme posterior features. Furthermore, we apply lattice based consensus decoding and system combination methods. In these methods, the effect of minimizing character instead of word errors is compared. The final system obtains a character error rate of 17.7% on the GALE 2006 evaluation data. Björn Hoffmeister, Christian Plahl, Peter Fritz, Georg Heigold, Jonas Lööf, Ralf Schlüter, Hermann Ney |
ASRU | 1 |
| 2007 | Cross-Site and Intra-Site ASR System Combination: Comparisons on Lattice and 1-Best MethodsabstractWe evaluate system combination techniques for automatic speech recognition using systems from multiple sites who participated in the TC-STAR 2006 evaluation. Both lattice and 1-best combination techniques are tested for cross-site and intra-site tasks. For pairwise combinations the lattice based approaches can outperform 1-best ROVER with confidence scores, but 1-best ROVER results are equal (or even better) when combining three or four systems. Björn Hoffmeister, Dustin Hillard, Stefan Hahn, Ralf Schlüter, Mari Ostendorf, Hermann Ney |
ICASSP (4) | 1 |
| 2007 | The RWTH 2007 TC-STAR evaluation system for european English and SpanishabstractIn this work, the RWTH automatic speech recognition systems developed for the third TC-STAR evaluation campaign 2007 are presented.The RWTH systems make systematic use of internal system combination, combining systems with differences in feature extraction, adaptation methods, and training data used.To take advantage of this, novel feature extraction methods were employed; this year saw the introduction of Gammatone features and MLP based phone posterior features.Further improvements were achieved using unsupervised training, and it is notable that these improvements were achieved using a fairly low amount of automatically transcribed data.Also contributing to the improvements over last year was the switch to MPE training, and the introduction of projecting SAT transforms. Jonas Lööf, Christian Gollan, Stefan Hahn, Georg Heigold, Björn Hoffmeister, Christian Plahl, David Rybach, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |
| 2006 | Frame based system combination and a comparison with weighted ROVER and CNCabstractIn this paper we present a novel ASR system combination technique able to combine systems producing word graphs of different structure and with different segmentations. The new method is based on the definition of a time frame-wise word error cost function in a minimum Bayes risk framework. In contrast to confusion network combination (CNC), it preserves both the word graph structure and the word boundaries. First experimental results are presented on the European Parliament Plenary Sessions (EPPS) task for European Spanish and British English. The new approach to system combination is compared to both ROVER and CNC. In addition, we also apply datadriven weighting schemes for all system combination approaches addressed in this work. For the experiments presented, a variety of internal systems as well as an additional external system were combined. Index Terms: speech recognition, system combination, word posteriors. 1. Björn Hoffmeister, Tobias Klein, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 1 |
| 2006 | The 2006 RWTH parliamentary speeches transcription systemabstractIn this work, investigations in the course of the developement of RWTH automatic speech recognition systems developed for the second TC-STAR evaluation campaign 2006 are presented.The systems were designed to transcribe parliamentary speeches taken from the European Parliament Plenary Sessions (EPPS) in European English and Spanish, as well as speeches from the Spanish Parliament.The RWTH systems apply a two pass search strategy with a fourgram one-pass decoder including a fast vocal tract length normalization variant as first pass.The systems further include several adaptation and normalization methods, minimum classification error trained models, and bayes risk minimization.For all relevant individual components contrastive results are presented on the EPPS Spanish and English data, including investigations which did not yet enter the evaluation systems. Jonas Lööf, Maximilian Bisani, Christian Gollan, Georg Heigold, Björn Hoffmeister, Christian Plahl, Ralf Schlüter, Hermann Ney |
INTERSPEECH | 5 |