EDBT 2026 Demo / reviewers in the wild / expert
Reinhold Häb-Umbach
dblp:88/6542 · also Reinhold Haeb, Reinhold Haeb-Umbach
· DBLP profile ↗
160ranked-venue papers
20as first author
27since 2021 · last 2025
0000-0001-9468-7330ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 137 · 14 first-author · 24 since 2021Artificial intelligence and machine learning · 83 · 11 first-author · 12 since 2021Computer networks · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 30+ Years of Source Separation Research: Achievements and Future ChallengesabstractSource separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50thanniversary, we review the major contributions and advancements in the past three decades in the speech, audio, and music SS research field. We will cover both single- and multi-channel SS approaches. We will also look back on key efforts to foster a culture of scientific evaluation in the research field, including challenges, performance metrics, and datasets. We will conclude by discussing current trends and future research directions. Shoko Araki, Nobutaka Ito, Reinhold Häb-Umbach, Gordon Wichern, Yuki Mitsufuji |
ICASSP | 3 |
| 2025 | Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture ModelsabstractWe propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (vMFMM) for diarization in a joint statistical framework. Through the integration, both spatial and spectral information are exploited for diarization and separation. We also develop a method for counting the number of active speakers in a segment of a meeting to support block-wise processing. While the total number of speakers in a meeting may be known, it is usually not known on a per-segment level. With the proposed speaker counting, joint diarization and source separation can be done segment-by-segment, and the permutation problem across segments is solved, thus allowing for block-online processing in the future. Experimental results on the LibriCSS meeting corpus show that the integrated approach outperforms a cascaded approach of diarization and speech enhancement in terms of WER, both on a per-segment and on a per-meeting level. Tobias Cord-Landwehr, Christoph Böddeker, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2025 | Speech Synthesis along Perceptual Voice Quality DimensionsabstractWhile expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities (PVQs) recognized by phonetic experts, which are speech properties at a lower level of abstraction. The ability to manipulate PVQs can be a valuable tool for teaching speech pathologists in training or voice actors. In this paper, we integrate a Conditional Continuous-Normalizing-Flow-based method into a Text-to-Speech system to modify perceptual voice attributes on a continuous scale. Unlike previous approaches, our system avoids direct manipulation of acoustic correlates and instead learns from examples. We demonstrate the system's capability by manipulating four voice qualities: Roughness, breathiness, resonance and weight. Phonetic experts evaluated these modifications, both for seen and unseen speaker conditions. The results highlight both the system's strengths and areas for improvement. Frederik Rautenberg, Michael Kuhlmann, Fritz Seebauer, Jana Wiechmann, Petra Wagner, Reinhold Häb-Umbach |
ICASSP | 6 |
| 2025 | Spatio-Spectral Diarization of Meetings by Combining TDOA-based Segmentation and Speaker Embedding-based ClusteringabstractWe propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data nor prior knowledge about the number or placement of microphones. It works for both a compact microphone array and distributed microphones, with minor adjustments. Due to its superior handling of overlapping speech during segmentation, the proposed pipeline significantly outperforms the single-channel pyannote approach, both in a scenario with a compact microphone array and in a setup with distributed microphones. Additionally, we show that, unlike fully spatial diarization pipelines, the proposed system can correctly track speakers when they change positions. Tobias Cord-Landwehr, Tobias Gburrek, Marc Deegen, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2025 | Towards Frame-level Quality Predictions of Synthetic SpeechabstractKuhlmann M, Seebauer FM, Wagner P, Haeb-Umbach R. Towards Frame-level Quality Predictions of Synthetic Speech. In: Interspeech 2025. Interspeech. Baixas: International Speech Communication Association; 2025: 2300-2304. Michael Kuhlmann, Fritz Seebauer, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2025 | Synthesizing Speech with Selected Perceptual Voice Qualities - A Case Study with Creaky VoiceabstractRautenberg F, Seebauer FM, Wiechmann J, Kuhlmann M, Wagner P, Haeb-Umbach R. Synthesizing Speech with Selected Perceptual Voice Qualities – A Case Study with Creaky Voice. In: Interspeech 2025. ISCA: ISCA; 2025: 1633-1637. Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann, Michael Kuhlmann, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 6 |
| 2024 | Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting ScenariosabstractWe propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions with two active speakers to lie on the shortest path on a sphere between the points given by the d-vectors of each of the active speakers. Using those frame-wise speaker embeddings in clustering-based diarization outperforms segment-level clustering-based diarization systems such as VBx and Spectral Clustering. By extending our approach to a mixture-model-based diarization, the performance can be further improved, approaching the diarization error rates of diarization systems that use a dedicated overlap detection, and outperforming these systems when also employing an additional overlap detection. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2024 | Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignmentabstractDiarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker.Particularly in the presence of overlapping or noisy speech, these systems have problems reliably assigning the correct speaker labels, leading to a significant amount of speaker confusion errors.We propose to add segment-level speaker reassignment to address this issue.By revisiting, after speech enhancement, the speaker attribution for each segment, speaker confusion errors from the initial diarization stage are significantly reduced.Through experiments across different system configurations and datasets, we further demonstrate the effectiveness and applicability in various domains.Our results show that segment-level speaker reassignment successfully rectifies at least 40% of speaker confusion word errors, highlighting its potential for enhancing diarization accuracy in meeting transcription systems. Christoph Böddeker, Tobias Cord-Landwehr, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2024 | Combining TF-GridNet And Mixture Encoder For Continuous Speech Separation For Meeting TranscriptionabstractMany real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement. Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Böddeker, Ralf Schlüter, Reinhold Häb-Umbach |
SLT | 6 |
| 2024 | TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker EmbeddingsabstractSince diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) diarization approach, which assumes that initial speaker embeddings are available. We replace the final combined speaker activity estimation network of TS-VAD with a network that produces speaker activity estimates at a time-frequency resolution. Those act as masks for source extraction, either via masking or via beamforming. The technique can be applied both for single-channel and multi-channel input and, in both cases, achieves a new state-of-the-art word error rate (WER) on the LibriCSS meeting data recognition task. We further compute speaker-aware and speaker-agnostic WERs to isolate the contribution of diarization errors to the overall WER performance. Christoph Böddeker, Aswin Shanmugam Subramanian, Gordon Wichern, Reinhold Häb-Umbach, Jonathan Le Roux |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting DiarizationabstractUsing a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2023 | On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition SystemsabstractWe propose a general framework to compute the word error rate (WER) of ASR systems that process recordings containing multiple speakers at their input and that produce multiple output word sequences (MIMO). Such ASR systems are typically required, e.g., for meeting transcription. We provide an efficient implementation based on a dynamic programming search in a multi-dimensional Levenshtein distance tensor under the constraint that a reference utterance must be matched consistently with one hypothesis output. This also results in an efficient implementation of the ORC WER which previously suffered from exponential complexity. We give an overview of commonly used WER definitions for multi-speaker scenarios and show that they are specializations of the above MIMO WER tuned to particular application scenarios. We conclude with a discussion of the pros and cons of the various WER definitions and a recommendation when to use which. Thilo von Neumann, Christoph Böddeker, Keisuke Kinoshita, Marc Delcroix, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2023 | Mixture Encoder for Joint Speech Separation and Recognition
Simon Berger, Peter Vieting, Christoph Böddeker, Ralf Schlüter, Reinhold Häb-Umbach |
INTERSPEECH | 5 |
| 2023 | A Teacher-Student Approach for Extracting Informative Speaker Embeddings From Speech MixturesabstractWe introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture.To allow for supervised training, a teacherstudent approach is employed: the teacher computes the target embeddings from each speaker's utterance before the utterances are added to form the mixture, and the student embedding extractor is then tasked to reproduce those embeddings from the speech mixture at its input.The system much more reliably verifies the presence or absence of a given speaker in a mixture than a conventional speaker embedding extractor, and even exhibits comparable performance to a multi-channel approach that exploits spatial information for embedding extraction.Further, it is shown that a speaker embedding computed from a mixture can be used to check for the presence of that speaker in another mixture. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
INTERSPEECH | 5 |
| 2023 | Segment-Less Continuous Speech Separation of Meetings: Training and Evaluation CriteriaabstractContinuous Speech Separation (CSS) has been proposed to address speech overlaps during the analysis of realistic meeting-like conversations by eliminating any overlaps before further processing. CSS separates a recording of arbitrarily many speakers into a small number of overlap-free output channels, where each output channel may contain speech of multiple speakers. Often, a separation model is trained with Utterance-level Permutation Invariant Training (uPIT), which exclusively maps a speaker to an output channel, and applied in a sliding window approach called stitching. Recently, we introduced an alternative training scheme called Graph-PIT that teaches the separator to produce a speaker-shared output channel format without stitching. It can handle an arbitrary number of speakers as long as the number of overlapping speakers is never larger than the number of output channels. Models trained in this way are able to perform segment-less CSS, i.e., without stitching, and achieve comparable and often better separation quality than the conventional CSS with uPIT and stitching. In this contribution, we further investigate the Graph-PIT training scheme. We show in extended experiments that Graph-PIT also works in challenging reverberant conditions. We simplify the training schedule for Graph-PIT with the recently proposed Source Aggregated Signal-to-Distortion Ratio (SA-SDR) loss, which eliminates unfavorable properties of the previously used A-SDR loss to enable training with Graph-PIT from scratch. Furthermore, we introduce novel signal-level evaluation metrics for meeting scenarios, namely the source-aggregated scale- and convolution-invariant Signal-to-Distortion Ratio (SA-SI-SDR and SA-CI-SDR), which are generalizations of the commonly used SDR-based metrics for the CSS case. Thilo von Neumann, Keisuke Kinoshita, Christoph Böddeker, Marc Delcroix, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Threshold Independent Evaluation of Sound Event Detection ScoresabstractPerforming an adequate evaluation of sound event detection (SED) systems is far from trivial and is still subject to ongoing research. The recently proposed polyphonic sound detection (PSD)-receiver operating characteristic (ROC) and PSD score (PSDS) make an important step into the direction of an evaluation of SED systems which is independent from a certain decision threshold. This allows to obtain a more complete picture of the overall system behavior which is less biased by threshold tuning. Yet, the PSD-ROC is currently only approximated using a finite set of thresholds. The choice of the thresholds used in approximation, however, can have a severe impact on the resulting PSDS. In this paper we propose a method which allows for computing system performance on an evaluation set for all possible thresholds jointly, enabling accurate computation not only of the PSD-ROC and PSDS but also of other collar-based and intersection-based performance curves. It further allows to select the threshold which best fulfills the requirements of a given application. Source code is publicly available in our SED evaluation package sed_scores_eval1. Janek Ebbers, Reinhold Häb-Umbach, Romain Serizel |
ICASSP | 2 |
| 2022 | On Synchronization of Wireless Acoustic Sensor Networks in the Presence of Time-Varying Sampling Rate Offsets and Speaker ChangesabstractA wireless acoustic sensor network records audio signals with sampling time and sampling rate offsets between the audio streams, if the analog-digital converters (ADCs) of the network devices are not synchronized. Here, we introduce a new sampling rate offset model to simulate time-varying sampling frequencies caused, for example, by temperature changes of ADC crystal oscillators, and propose an estimation algorithm to handle this dynamic aspect in combination with changing acoustic source positions. Furthermore, we show how estimates of the distances between microphones and human speakers can be used to determine the sampling time offsets. This enables a synchronization of the audio streams to reflect the physical time differences of flight. Tobias Gburrek, Joerg Schmalenstroeer, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2022 | SA-SDR: A Novel Loss Function for Separation of Meeting Style DataabstractMany state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal perfectly or if the reference signal contains silence, e.g., when a two-output separator processes a single-speaker recording. Many modifications to the plain SDR have been proposed that trade-off between making the loss more robust and distorting its value. We propose to switch from a mean over the SDRs of each individual output channel to a global SDR over all output channels at the same time, which we call source-aggregated SDR (SA-SDR). This makes the loss robust against silence and perfect reconstruction as long as at least one reference signal is not silent. We experimentally show that our proposed SA-SDR is more stable and preferable over other well-known modifications when processing meeting-style data that typically contains many silent or single-speaker regions. Thilo von Neumann, Keisuke Kinoshita, Christoph Böddeker, Marc Delcroix, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2022 | An Initialization Scheme for Meeting Separation with Spatial Mixture Models
Christoph Böddeker, Tobias Cord-Landwehr, Thilo von Neumann, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2022 | Utterance-by-utterance overlap-aware neural diarization with Graph-PITabstractRecent speaker diarization studies showed that integration of end-to-end neural diarization (EEND) and clustering-based diarization is a promising approach for achieving state-of-the-art performance on various tasks.Such an approach first divides an observed signal into fixed-length segments, then performs segment-level local diarization based on an EEND module, and merges the segment-level results via clustering to form a final global diarization result.The segmentation is done to limit the number of speakers in each segment since the current EEND cannot handle a large number of speakers.In this paper, we argue that such an approach involving the segmentation has several issues; for example, it inevitably faces a dilemma that larger segment sizes increase both the context available for enhancing the performance and the number of speakers for the local EEND module to handle.To resolve such a problem, this paper proposes a novel framework that performs diarization without segmentation.However, it can still handle challenging data containing many speakers and a significant amount of overlapping speech.The proposed method can take an entire meeting for inference and perform utterance-by-utterance diarization that clusters utterance activities in terms of speakers.To this end, we leverage a neural network training scheme called Graph-PIT proposed recently for neural source separation.Experiments with simulated active-meeting-like data and CALLHOME data show the superiority of the proposed approach over the conventional methods. Keisuke Kinoshita, Thilo von Neumann, Marc Delcroix, Christoph Böddeker, Reinhold Häb-Umbach |
INTERSPEECH | 5 |
| 2022 | Investigation into Target Speaking Rate Adaptation for Voice ConversionabstractDisentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained with non-parallel and unlabeled speech data.However, previous approaches perform disentanglement only implicitly via some sort of information bottleneck or normalization, where it is usually hard to find a good trade-off between voice conversion and content reconstruction.Further, previous works usually do not consider an adaptation of the speaking rate to the target speaker or they put some major restrictions to the data or use case.Therefore, the contribution of this work is two-fold.First, we employ an explicit and fully unsupervised disentanglement approach, which has previously only been used for representation learning, and show that it allows to obtain both superior voice conversion and content reconstruction.Second, we investigate simple and generic approaches to linearly adapt the length of a speech signal, and hence the speaking rate, to a target speaker and show that the proposed adaptation allows to increase the speaking rate similarity with respect to the target speaker. Michael Kuhlmann, Fritz Seebauer, Janek Ebbers, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 5 |
| 2021 | Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech SeparationabstractTime-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine neural network supported multi-channel source separation with a time-domain training objective function. For the objective we propose to use a convolutive transfer function invariant Signal-to-Distortion Ratio (CI-SDR) based loss. While this is a well-known evaluation metric (BSS Eval), it has not been used as a training objective before. To show the effectiveness, we demonstrate the performance on LibriSpeech based reverberant mixtures. On this task, the proposed system approaches the error rate obtained on single-source non-reverberant input, i.e., LibriSpeech test clean, with a difference of only 1.2 percentage points, thus outperforming a conventional permutation invariant training based system and alternative objectives like Scale Invariant Signal-to-Distortion Ratio by a large margin. Christoph Böddeker, Wangyou Zhang, Tomohiro Nakatani, Keisuke Kinoshita, Tsubasa Ochiai, Marc Delcroix, Naoyuki Kamo, Yanmin Qian, Reinhold Häb-Umbach |
ICASSP | 9 |
| 2021 | Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech RepresentationsabstractIn this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentanglement method does neither need parallel data nor any supervision. We show that the proposed technique is capable of separating speaker and content traits into the two different representations and show competitive speaker-content disentanglement performance compared to other unsupervised approaches. We further demonstrate an increased robustness of the content representation against a train-test mismatch compared to spectral features, when used for phone recognition. Janek Ebbers, Michael Kuhlmann, Tobias Cord-Landwehr, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2021 | Iterative Geometry Calibration from Distance Estimates for Wireless Acoustic Sensor NetworksabstractIn this paper we present an approach to geometry calibration in wireless acoustic sensor networks, whose nodes are assumed to be equipped with a compact microphone array. The proposed approach solely works with estimates of the distances between acoustic sources and the nodes that record these sources. It consists of an iterative weighted least squares localization procedure, which is initialized by multidimensional scaling. Alongside the sensor node locations, also the positions of the acoustic sources are estimated. Furthermore, we derive the Cramer-Rao lower bound (CRLB) for source and sensor position estimation, and show by simulation that the estimator is efficient. Tobias Gburrek, Joerg Schmalenstroeer, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2021 | End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced FrontendabstractRecently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performance gap between anechoic and reverberant conditions. In this work, we focus on the multichannel multi-speaker reverberant condition, and propose to extend our previous framework for end-to-end dereverberation, beamforming, and speech recognition with improved numerical stability and advanced frontend subnetworks including voice activity detection like masks. The techniques significantly stabilize the end-to-end training process. The experiments on the spatialized wsj1-2mix corpus show that the proposed system achieves about 35% WER relative reduction compared to our conventional multi-channel E2E ASR system, and also obtains decent speech dereverberation and separation performance (SDR=12.5 dB) in the reverberant multi-speaker condition while trained only with the ASR criterion. Wangyou Zhang, Christoph Böddeker, Shinji Watanabe 0001, Tomohiro Nakatani, Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Naoyuki Kamo, Reinhold Häb-Umbach, Yanmin Qian |
ICASSP | 9 |
| 2021 | Graph-PIT: Generalized Permutation Invariant Training for Continuous Separation of Arbitrary Numbers of SpeakersabstractAutomatic transcription of meetings requires handling of overlapped speech, which calls for continuous speech separation (CSS) systems. The uPIT criterion was proposed for utterance-level separation with neural networks and introduces the constraint that the total number of speakers must not exceed the number of output channels. When processing meeting-like data in a segment-wise manner, i.e., by separating overlapping segments independently and stitching adjacent segments to continuous output streams, this constraint has to be fulfilled for any segment. In this contribution, we show that this constraint can be significantly relaxed. We propose a novel graph-based PIT criterion, which casts the assignment of utterances to output channels in a graph coloring problem. It only requires that the number of concurrently active speakers must not exceed the number of output channels. As a consequence, the system can process an arbitrary number of speakers and arbitrarily long segments and thus can handle more diverse scenarios. Further, the stitching algorithm for obtaining a consistent output order in neighboring segments is of less importance and can even be eliminated completely, not the least reducing the computational effort. Experiments on meeting-style WSJ data show improvements in recognition performance over using the uPIT criterion. Thilo von Neumann, Keisuke Kinoshita, Christoph Böddeker, Marc Delcroix, Reinhold Häb-Umbach |
Interspeech | 5 |
| 2021 | Far-Field Automatic Speech RecognitionabstractThe machine recognition of speech spoken at a distance from the microphones, known as far-field automatic speech recognition (ASR), has received a significant increase in attention in science and industry, which caused or was caused by an equally significant improvement in recognition accuracy. Meanwhile, it has entered the consumer market with digital home assistants with a spoken language interface being its most prominent application. Speech recorded at a distance is affected by various acoustic distortions, and consequently, quite different processing pipelines have emerged compared with ASR for close-talk speech. A signal enhancement front end for dereverberation, source separation, and acoustic beamforming is employed to clean up the speech, and the back-end ASR engine is robustified by multicondition training and adaptation. We will also describe the so-called end-to-end approach to ASR, which is a new promising architecture that has recently been extended to the far-field scenario. This tutorial article gives an account of the algorithms used to enable accurate speech recognition from a distance, and it will be seen that, although deep learning has a significant share in the technological breakthroughs, a clever combination with traditional signal processing can lead to surprisingly effective solutions. Reinhold Häb-Umbach, Jahn Heymann, Lukas Drude, Shinji Watanabe 0001, Marc Delcroix, Tomohiro Nakatani |
Proc. IEEE | 1 |
| 2020 | Jointly Optimal Dereverberation and BeamformingabstractWe previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a Weighted Prediction Error (WPE) dereverberation filter and a conventional Minimum-Power Distortionless Response (MPDR) beamformer. However, it has not been fully investigated which components in the convolutional beamformer yield such superiority. To this end, this paper presents a new derivation of the convolutional beamformer that allows us to factorize it into a WPE dereverberation filter, and a special type of a (non-convolutional) beamformer, referred to as a weighted MPDR (wM-PDR) beamformer, without loss of optimality. With experiments, we show that the superiority of the convolutional beamformer in fact comes from its wMPDR part. Christoph Böddeker, Tomohiro Nakatani, Keisuke Kinoshita, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2020 | Demystifying TasNet: A Dissecting ApproachabstractIn recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the gains of the time-domain audio separation network (TasNet) approach by gradually replacing components of an utterance-level permutation invariant training (u-PIT) based separation system in the frequency domain until the TasNet system is reached, thus blending components of frequency domain approaches with those of time domain approaches. Some of the intermediate variants achieve comparable signal-to-distortion ratio (SDR) gains to TasNet, but retain the advantage of frequency domain processing: compatibility with classic signal processing tools such as frequency-domain beamforming and the human interpretability of the masks. Furthermore, we show that the scale invariant signal-to-distortion ratio (si-SDR) criterion used as loss function in TasNet is related to a logarithmic mean square error criterion and that it is this criterion which contributes most reliable to the performance advantage of TasNet. Finally, we critically assess which gains in a noise-free single channel environment generalize to more realistic reverberant conditions. Jens Heitkaemper, Darius Jakobeit, Christoph Böddeker, Lukas Drude, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2020 | End-to-End Training of Time Domain Audio Separation and RecognitionabstractThe rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recognition. We here demonstrate how to combine a separation module based on a Convolutional Time domain Audio Separation Network (Conv-TasNet) with an E2E speech recognizer and how to train such a model jointly by distributing it over multiple GPUs or by approximating truncated back-propagation for the convolutional front-end. To put this work into perspective and illustrate the complexity of the design space, we provide a compact overview of single-channel multi-speaker recognition systems. Our experiments show a word error rate of 11.0% on WSJ0-2mix and indicate that our joint time domain model can yield substantial improvements over cascade DNN-HMM and monolithic E2E frequency domain systems proposed so far. Thilo von Neumann, Keisuke Kinoshita, Lukas Drude, Christoph Böddeker, Marc Delcroix, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 7 |
| 2020 | Statistical and Neural Network Based Speech Activity Detection in Non-Stationary Acoustic EnvironmentsabstractSpeech activity detection (SAD), which often rests on the fact that the noise is "more" stationary than speech, is particularly challenging in non-stationary environments, because the time variance of the acoustic scene makes it difficult to discriminate speech from noise. We propose two approaches to SAD, where one is based on statistical signal processing, while the other utilizes neural networks. The former employes sophisticated signal processing to track the noise and speech energies and is meant to support the case for a resource efficient, unsupervised signal processing approach. The latter introduces a recurrent network layer that operates on short segments of the input speech to do temporal smoothing in the presence of non-stationary noise. The systems are tested on the Fearless Steps challenge, which consists of the transmission data from the Apollo-11 space mission. The statistical SAD achieves comparable detection performance to earlier proposed neural network based SADs, while the neural network based approach leads to a decision cost function of 1.07% on the evaluation set of the 2020 Fearless Steps Challenge, which sets a new state of the art. Jens Heitkaemper, Joerg Schmalenstroeer, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2020 | Multi-Path RNN for Hierarchical Modeling of Long Sequential Data and its Application to Speaker Stream SeparationabstractRecently, the source separation performance was greatly improved by time-domain audio source separation based on dualpath recurrent neural network (DPRNN).DPRNN is a simple but effective model for a long sequential data.While DPRNN is quite efficient in modeling a sequential data of the length of an utterance, i.e., about 5 to 10 second data, it is harder to apply it to longer sequences such as whole conversations consisting of multiple utterances.It is simply because, in such a case, the number of time steps consumed by its internal module called inter-chunk RNN becomes extremely large.To mitigate this problem, this paper proposes a multi-path RNN (MPRNN), a generalized version of DPRNN, that models the input data in a hierarchical manner.In the MPRNN framework, the input data is represented at several (≥ 3) time-resolutions, each of which is modeled by a specific RNN sub-module.For example, the RNN sub-module that deals with the finest resolution may model temporal relationship only within a phoneme, while the RNN sub-module handling the most coarse resolution may capture only the relationship between utterances such as speaker information.We perform experiments using simulated dialogue-like mixtures and show that MPRNN has greater model capacity, and it outperforms the current state-of-the-art DPRNN framework especially in online processing scenarios. Keisuke Kinoshita, Thilo von Neumann, Marc Delcroix, Tomohiro Nakatani, Reinhold Häb-Umbach |
INTERSPEECH | 5 |
| 2020 | Multi-Talker ASR for an Unknown Number of Sources: Joint Training of Source Counting, Separation and ASRabstractMost approaches to multi-talker overlapped speech separation and recognition assume that the number of simultaneously active speakers is given, but in realistic situations, it is typically unknown. To cope with this, we extend an iterative speech extraction system with mechanisms to count the number of sources and combine it with a single-talker speech recognizer to form the first end-to-end multi-talker automatic speech recognition system for an unknown number of active speakers. Our experiments show very promising performance in counting accuracy, source separation and speech recognition on simulated clean mixtures from WSJ0-2mix and WSJ0-3mix. Among others, we set a new state-of-the-art word error rate on the WSJ0-2mix database. Furthermore, our system generalizes well to a larger number of speakers than it ever saw during training, as shown in experiments with the WSJ0-4mix database. Thilo von Neumann, Christoph Böddeker, Lukas Drude, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Reinhold Häb-Umbach |
INTERSPEECH | 7 |
| 2020 | Jointly Optimal Denoising, Dereverberation, and Source SeparationabstractThis article proposes methods that can optimize a Convolutional BeamFormer (CBF) for jointly performing denoising, dereverberation, and source separation (DN+DR+SS) in a computationally efficient way. Conventionally, a cascade configuration, composed of a Weighted Prediction Error minimization (WPE) dereverberation filter followed by a Minimum Variance Distortionless Response (MVDR) beamformer, has been used as the state-of-the-art frontend of far-field speech recognition, even though this approach's overall optimality is not guaranteed. In the blind signal processing area, an approach for jointly optimizing dereverberation and source separation (DR+SS) has been proposed; however, it requires huge computing cost, and has not been extended for applications to DN+DR+SS. To overcome the above limitations, this paper develops new approaches for jointly optimizing DN+DR+SS in a computationally much more efficient way. To this end, we first present an objective function to optimize a CBF for performing DN+DR+SS based on maximum likelihood estimation on an assumption that the steering vectors of the target signals are given or can be estimated, e.g., using a neural network. This paper refers to a CBF optimized by this objective function as a weighted Minimum-Power Distortionless Response (wMPDR) CBF. Then, we derive two algorithms for optimizing a wMPDR CBF based on two different ways of factorizing a CBF into WPE filters and beamformers: one based on an extension of the conventional joint optimization approach proposed for DR+SS and another based on a novel technique. Experiments using noisy reverberant sound mixtures show that the proposed optimization approaches greatly improve the performance of the speech enhancement in comparison with the conventional cascade configuration in terms of signal distortion measures and ASR performance. The proposed approaches also greatly reduce the computing cost with improved estimation accuracy in comparison with the conventional joint optimization approach. Tomohiro Nakatani, Christoph Böddeker, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2019 | An Investigation into the Effectiveness of Enhancement in ASR Training and Test for Chime-5 Dinner Party TranscriptionabstractDespite the strong modeling power of neural network acoustic models, speech enhancement has been shown to deliver additional word error rate improvements if multi-channel data is available. However, there has been a longstanding debate whether enhancement should also be carried out on the ASR training data. In an extensive experimental evaluation on the acoustically very challenging CHiME-5 dinner party data we show that: (i) cleaning up the training data can lead to substantial error rate reductions, and (ii) enhancement in training is advisable as long as enhancement in test is at least as strong as in training. This approach stands in contrast and delivers larger gains than the common strategy reported in the literature to augment the training database with additional artificially degraded speech. Together with an acoustic model topology consisting of initial CNN layers followed by factorized TDNN layers we achieve with 41.6 % and 43.2 % WER on the DEV and EVAL test sets, respectively, a new single-system state-of-the-art result on the CHiME-5 data. This is a 8 % relative improvement compared to the best word error rate published so far for a speech recognizer without system combination. Catalin Zorila, Christoph Böddeker, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ASRU | 4 |
| 2019 | Unsupervised Training of a Deep Clustering Model for Multichannel Blind Source SeparationabstractWe propose a training scheme to train neural network-based source separation algorithms from scratch when parallel clean data is unavailable. In particular, we demonstrate that an unsupervised spatial clustering algorithm is sufficient to guide the training of a deep clustering system. We argue that previous work on deep clustering requires strong supervision and elaborate on why this is a limitation. We demonstrate that (a) the single-channel deep clustering system trained according to the proposed scheme alone is able to achieve a similar performance as the multi-channel teacher in terms of word error rates and (b) initializing the spatial clustering approach with the deep clustering result yields a relative word error rate reduction of 26% over the unsupervised teacher. Lukas Drude, Daniel Hasenklever, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2019 | Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASRabstractSignal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore enabled its use in online applications. For this algorithm, the estimation of the power spectral density (PSD) of the anechoic signal plays an important role and strongly influences its performance. Recently, we showed that using a neural network PSD estimator leads to improved performance for online automatic speech recognition. This, however, comes at a price. To train the network, we require parallel data, i.e., utterances simultaneously available in clean and reverberated form. Here we propose to overcome this limitation by training the network jointly with the acoustic model of the speech recognizer. To be specific, the gradients computed from the cross-entropy loss between the target senone sequence and the acoustic model network output is backpropagated through the complex-valued dereverberation filter estimation to the neural network for PSD estimation. Evaluation on two databases demonstrates improved performance for on-line processing scenarios while imposing fewer requirements on the available training data and thus widening the range of applications. Jahn Heymann, Lukas Drude, Reinhold Häb-Umbach, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 3 |
| 2019 | All-neural Online Source Separation, Counting, and Diarization for Meeting AnalysisabstractAutomatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant progress has been made on individual tasks, this paper presents for the first time an all-neural approach to simultaneous speaker counting, diarization and source separation. The NN-based estimator operates in a block-online fashion and tracks speakers even if they remain silent for a number of time blocks, thus learning a stable output order for the separated sources. The neural network is recurrent over time as well as over the number of sources. The simulation experiments show that state of the art separation performance is achieved, while at the same time delivering good diarization and source counting results. It even generalizes well to an unseen large number of blocks. Thilo von Neumann, Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 6 |
| 2019 | Unsupervised Training of Neural Mask-Based BeamformingabstractWe present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood criterion derived from a spatial mixture model of the observations. It is trained from scratch without requiring any parallel data consisting of degraded input and clean training targets. Thus, training can be carried out on real recordings of noisy speech rather than simulated ones. In contrast to previous work on unsupervised training of neural mask estimators, our approach avoids the need for a possibly pre-trained teacher model entirely. We demonstrate the effectiveness of our approach by speech recognition experiments on two different datasets: one mainly deteriorated by noise (CHiME 4) and one by reverberation (REVERB). The results show that the performance of the proposed system is on par with a supervised system using oracle target masks for training and with a system trained using a model-based teacher. Lukas Drude, Jahn Heymann, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2019 | Guided Source Separation Meets a Strong ASR Backend: Hitachi/Paderborn University Joint Investigation for Dinner Party ASRabstractIn this paper, we present Hitachi and Paderborn University's joint effort for automatic speech recognition (ASR) in a dinner party scenario.The main challenges of ASR systems for dinner party recordings obtained by multiple microphone arrays are (1) heavy speech overlaps, (2) severe noise and reverberation, (3) very natural conversational content, and possibly (4) insufficient training data.As an example of a dinner party scenario, we have chosen the data presented during the CHiME-5 speech recognition challenge, where the baseline ASR had a 73.3% word error rate (WER), and even the best performing system at the CHiME-5 challenge had a 46.1% WER.We extensively investigated a combination of the guided source separation-based speech enhancement technique and an already proposed strong ASR backend and found that a tight combination of these techniques provided substantial accuracy improvements.Our final system achieved WERs of 39.94% and 41.64% for the development and evaluation data, respectively, both of which are the best published results for the dataset.We also investigated with additional training data on the official small data in the CHiME-5 corpus to assess the intrinsic difficulty of this ASR task. Naoyuki Kanda, Christoph Böddeker, Jens Heitkaemper, Yusuke Fujita, Shota Horiguchi, Kenji Nagamatsu, Reinhold Häb-Umbach |
INTERSPEECH | 7 |
| 2019 | Multi-Channel Block-Online Source Extraction Based on Utterance AdaptationabstractThis paper deals with multi-channel speech recognition in scenarios with multiple speakers. Recently, the spectral characteristics of a target speaker, extracted from an adaptation utterance, have been used to guide a neural network mask estimator to focus on that speaker. In this work we present two variants of speakeraware neural networks, which exploit both spectral and spatial information to allow better discrimination between target and interfering speakers. Thus, we introduce either a spatial preprocessing prior to the mask estimation or a spatial plus spectral speaker characterization block whose output is directly fed into the neural mask estimator. The target speaker’s spectral and spatial signature is extracted from an adaptation utterance recorded at the beginning of a session. We further adapt the architecture for low-latency processing by means of block-online beamforming that recursively updates the signal statistics. Experimental results show that the additional spatial information clearly improves source extraction, in particular in the same-gender case, and that our proposal achieves state-of-the-art performance in terms of distortion reduction and recognition accuracy. Juan M. Martín-Doñas, Jens Heitkaemper, Reinhold Häb-Umbach, Ángel M. Gómez, Antonio M. Peinado |
INTERSPEECH | 3 |
| 2019 | Privacy-Preserving Variational Information Feature Extraction for Domestic Activity Monitoring versus Speaker IdentificationabstractIn this paper we highlight the privacy risks entailed in deep neural network feature extraction for domestic activity monitoring. We employ the baseline system proposed in the Task 5 of the DCASE 2018 challenge and simulate a feature interception attack by an eavesdropper who wants to perform speaker identification. We then propose to reduce the aforementioned privacy risks by introducing a variational information feature extraction scheme that allows for good activity monitoring performance while at the same time minimizing the information of the feature representation, thus restricting speaker identification attempts. We analyze the resulting model’s composite loss function and the budget scaling factor used to control the balance between the performance of the trusted and attacker tasks. It is empirically demonstrated that the proposed method reduces speaker identification privacy risks without significantly deprecating the performance of domestic activity monitoring tasks. Alexandru Nelus, Janek Ebbers, Reinhold Häb-Umbach, Rainer Martin 0001 |
INTERSPEECH | 3 |
| 2018 | Exploring Practical Aspects of Neural Mask-Based Beamforming for Far-Field Speech RecognitionabstractThis work examines acoustic beamformers employing neural networks (NNs) for mask prediction as front -end for automatic speech recognition (ASR) systems for practical scenarios like voice-enabled home devices. To test the versatility of the mask predicting network, the system is evaluated with different recording hardware, different microphone array designs, and different acoustic models of the downstream ASR system. Significant gains in recognition accuracy are obtained in all configurations despite the fact that the NN had been trained on mismatched data. Unlike previous work, the NN is trained on a feature level objective, which gives some performance advantage over a mask related criterion. Furthermore, different approaches for realizing online, or adaptive, NN-based beamforming are explored, where the online algorithms still show significant gains compared to the baseline performance. Christoph Böddeker, Hakan Erdogan, Takuya Yoshioka, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2018 | Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source SeparationabstractDeep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on blocks yields a block permutation problem in that the index of each speaker may change between blocks. We here propose the joint modeling of spatial and spectral features to solve the block permutation problem and generalize DANs to multi-channel meeting recordings: The DAN acts as a spectral feature extractor for a subsequent model-based clustering approach. We first analyze different joint models in batch-processing scenarios and finally propose a block-online blind source separation algorithm. The efficacy of the proposed models is demonstrated on reverberant mixtures corrupted by real recordings of multi-channel background noise. We demonstrate that both the proposed batch-processing and the proposed block-online system outperform (a) a spatial-only model with a state-of-the-art frequency permutation solver and (b) a spectral-only model with an oracle block permutation solver in terms of signal to distortion ratio (SDR) gains. Lukas Drude, Takuya Higuchi, Keisuke Kinoshita, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2018 | Deep Attractor Networks for Speaker Re-Identification and Blind Source SeparationabstractDeep clustering (DC) and deep attractor networks (DANs) are a data-driven way to monaural blind source separation. Both approaches provide astonishing single channel performance but have not yet been generalized to block-online processing. When separating speech in a continuous stream with a block-online algorithm, it needs to be determined in each block which of the output streams belongs to whom. In this contribution we solve this block permutation problem by introducing an additional speaker identification embedding to the DAN model structure. We motivate this model decision by analyzing the embedding topology of DC and DANs and show, that DC and DANs themselves are not sufficient for speaker identification. This model structure (a) improves the signal to distortion ratio (SDR) over a DAN baseline and (b) provides up to 61% and up to 34% relative reduction in permutation error rate and re-identification error rate compared to an i-vector baseline, respectively. Lukas Drude, Thilo von Neumann, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2018 | Integrating Neural Network Based Beamforming and Weighted Prediction Error DereverberationabstractThe weighted prediction error (WPE) algorithm has proven to be a very successful dereverberation method for the REVERB challenge. Likewise, neural network based mask estimation for beamforming demonstrated very good noise suppression in the CHiME 3 and CHiME 4 challenges. Recently, it has been shown that this estimator can also be trained to perform dereverberation and denoising jointly. However, up to now a comparison of a neural beamformer and WPE is still missing, so is an investigation into a combination of the two. Therefore, we here provide an extensive evaluation of both and consequently propose variants to integrate deep neural network based beamforming with WPE. For these integrated variants we identify a consistent word error rate (WER) reduction on two distinct databases. In particular, our study shows that deep learning based beamforming benefits from a model-based dereverberation technique (i.e. WPE) and vice versa. Our key findings are: (a) Neural beamforming yields the lower WERs in comparison to WPE the more channels and noise are present. (b) Integration of WPE and a neural beamformer consistently outperforms all stand-alone systems. Lukas Drude, Christoph Böddeker, Jahn Heymann, Reinhold Häb-Umbach, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2018 | Full Bayesian Hidden Markov Model Variational Autoencoder for Acoustic Unit DiscoveryabstractThe invention of the Variational Autoencoder enables the application of Neural Networks to a wide range of tasks in unsupervised learning, including the field of Acoustic Unit Discovery (AUD). The recently proposed Hidden Markov Model Variational Autoencoder (HMMVAE) allows a joint training of a neural network based feature extractor and a structured prior for the latent space given by a Hidden Markov Model. It has been shown that the HMMVAE significantly outperforms pure GMM-HMM based systems on the AUD task. However, the HMMVAE cannot autonomously infer the number of acoustic units and thus relies on the GMM-HMM system for initialization. This paper introduces the Bayesian Hidden Markov Model Variational Autoencoder (BHMMVAE) which solves these issues by embedding the HMMVAE in a Bayesian framework with a Dirichlet Process Prior for the distribution of the acoustic units, and diagonal or full-covariance Gaussians as emission distributions. Experiments on TIMIT and Xitsonga show that the BHMMVAE is able to autonomously infer a reasonable number of acoustic units, can be initialized without supervision by a GMM-HMM system, achieves computationally efficient stochastic variational inference by using natural gradient descent, and, additionally, improves the AUD performance over the HMMVAE. Thomas Glarner, Patrick Hanebrink, Janek Ebbers, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2018 | Machine learning techniques for semantic analysis of dysarthric speech: An experimental study
Vladimir Despotovic, Oliver Walter, Reinhold Häb-Umbach |
Speech Commun. | 3 |
| 2017 | Optimizing neural-network supported acoustic beamforming by algorithmic differentiationabstractIn this paper we show how a neural network for spectral mask estimation for an acoustic beamformer can be optimized by algorithmic differentiation. Using the beamformer output SNR as the objective function to maximize, the gradient is propagated through the beamformer all the way to the neural network which provides the clean speech and noise masks from which the beamformer coefficients are estimated by eigenvalue decomposition. A key theoretical result is the derivative of an eigenvalue problem involving complex-valued eigenvectors. Experimental results on the CHiME-3 challenge database demonstrate the effectiveness of the approach. The tools developed in this paper are a key component for an end-to-end optimization of speech enhancement and speech recognition. Christoph Böddeker, Patrick Hanebrink, Lukas Drude, Jahn Heymann, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2017 | A generalized log-spectral amplitude estimator for single-channel speech enhancementabstractThe benefits of both a logarithmic spectral amplitude (LSA) estimation and a modeling in a generalized spectral domain (where short-time amplitudes are raised to a generalized power exponent, not restricted to magnitude or power spectrum) are combined in this contribution to achieve a better tradeoff between speech quality and noise suppression in single-channel speech enhancement. A novel gain function is derived to enhance the logarithmic generalized spectral amplitudes of noisy speech. Experiments on the CHiME-3 dataset show that it outperforms the famous minimum mean squared error (MMSE) LSA gain function of Ephraim and Malah in terms of noise suppression by 1.4 dB, while the good speech quality of the MMSE-LSA estimator is maintained. Aleksej Chinaev, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2017 | Beamnet: End-to-end training of a beamformer-supported multi-channel ASR systemabstractThis paper presents an end-to-end training approach for a beamformer-supported multi-channel ASR system. A neural network which estimates masks for a statistically optimum beamformer is jointly trained with a network for acoustic modeling. To update its parameters, we propagate the gradients from the acoustic model all the way through feature extraction and the complex valued beamforming operation. Besides avoiding a mismatch between the front-end and the back-end, this approach also eliminates the need for stereo data, i.e., the parallel availability of clean and noisy versions of the signals. Instead, it can be trained with real noisy multi-channel data only. Also, relying on the signal statistics for beamforming, the approach makes no assumptions on the configuration of the microphone array. We further observe a performance gain through joint training in terms of word error rate in an evaluation of the system on the CHiME 4 dataset. Jahn Heymann, Lukas Drude, Christoph Böddeker, Patrick Hanebrink, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2017 | Tight Integration of Spatial and Spectral Features for BSS with Deep Clustering EmbeddingsabstractRecent advances in discriminatively trained mask estimation networks to extract a single source utilizing beamforming techniques demonstrate, that the integration of statistical models and deep neural networks (DNNs) are a promising approach for robust automatic speech recognition (ASR) applications. In this contribution we demonstrate how discriminatively trained embeddings on spectral features can be tightly integrated into statistical model-based source separation to separate and transcribe overlapping speech. Good generalization to unseen spatial configurations is achieved by estimating a statistical model at test time, while still leveraging discriminative training of deep clustering embeddings on a separate training set. We formulate an expectation maximization (EM) algorithm which jointly estimates a model for deep clustering embeddings and complex-valued spatial observations in the short time Fourier transform (STFT) domain at test time. Extensive simulations confirm, that the integrated model outperforms (a) a deep clustering model with a subsequent beamforming step and (b) an EM-based model with a beamforming step alone in terms of signal to distortion ratio (SDR) and perceptually motivated metric (PESQ) gains. ASR results on a reverberated dataset further show, that the aforementioned gains translate to reduced word error rates (WERs) even in reverberant environments. Lukas Drude, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2017 | Hidden Markov Model Variational Autoencoder for Acoustic Unit DiscoveryabstractVariational Autoencoders (VAEs) have been shown to provide efficient neural-network-based approximate Bayesian inference for observation models for which exact inference is intractable. Its extension, the so-called Structured VAE (SVAE) allows inference in the presence of both discrete and continuous latent variables. Inspired by this extension, we developed a VAE with Hidden Markov Models (HMMs) as latent models. We applied the resulting HMM-VAE to the task of acoustic unit discovery in a zero resource scenario. Starting from an initial model based on variational inference in an HMM with Gaussian Mixture Model (GMM) emission probabilities, the accuracy of the acoustic unit discovery could be significantly improved by the HMM-VAE. In doing so we were able to demonstrate for an unsupervised learning task what is well-known in the supervised learning case: Neural networks provide superior modeling power compared to GMMs. Janek Ebbers, Jahn Heymann, Lukas Drude, Thomas Glarner, Reinhold Häb-Umbach, Bhiksha Raj |
INTERSPEECH | 5 |
| 2017 | Leveraging Text Data for Word Segmentation for Underresourced LanguagesabstractIn this contribution we show how to exploit text data to support word discovery from audio input in an underresourced target language. Given audio, of which a certain amount is transcribed at the word level, and additional unrelated text data, the approach is able to learn a probabilistic mapping from acoustic units to characters and utilize it to segment the audio data into words without the need of a pronunciation dictionary. This is achieved by three components: an unsupervised acoustic unit discovery system, a supervisedly trained acoustic unit-to-grapheme converter, and a word discovery system, which is initialized with a language model trained on the text data. Experiments for multiple setups show that the initialization of the language model with text data improves the word segementation performance by a large margin. Thomas Glarner, Benedikt T. Boenninghoff, Oliver Walter, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2017 | A study on transfer learning for acoustic event detection in a real life scenarioabstractIn this work, we address the limited availability of large annotated databases for real-life audio event detection by utilizing the concept of transfer learning. This technique aims to transfer knowledge from a source domain to a target domain, even if source and target have different feature distributions and label sets. We hypothesize that all acoustic events share the same inventory of basic acoustic building blocks and differ only in the temporal order of these acoustic units. We then construct a deep neural network with convolutional layers for extracting the acoustic units and a recurrent layer for capturing the temporal order. Under the above hypothesis, transfer learning from a source to a target domain with a different acoustic event inventory is realized by transferring the convolutional layers from the source to the target domain. The recurrent layer is, however, learnt directly from the target domain. Experiments on the transfer from a synthetic source database to the real-life target database of DCASE 2016 demonstrate that transfer learning leads to improved detection performance on average. However, the successful transfer to detect events which are very different from what was seen in the source domain, could not be verified. Prerna Arora, Reinhold Häb-Umbach |
MMSP | 2 |
| 2017 | Multi-stage coherence drift based sampling rate synchronization for acoustic beamformingabstractMulti-channel speech enhancement algorithms rely on a synchronous sampling of the microphone signals. This, however, cannot always be guaranteed, especially if the sensors are distributed in an environment. To avoid performance degradation the sampling rate offset needs to be estimated and compensated for. In this contribution we extend the recently proposed coherence drift based method in two important directions. First, the increasing phase shift in the short-time Fourier transform domain is estimated from the coherence drift in a Matched Filter-like fashion, where intermediate estimates are weighted by their instantaneous SNR. Second, an observed bias is removed by iterating between offset estimation and compensation by resampling a couple of times. The effectiveness of the proposed method is demonstrated by speech recognition results on the output of a beamformer with and without sampling rate offset compensation between the input channels. We compare MVDR and maximum-SNR beamformers in reverberant environments and further show that both benefit from a novel phase normalization, which we also propose in this contribution. Joerg Schmalenstroeer, Jahn Heymann, Lukas Drude, Christoph Böddeker, Reinhold Häb-Umbach |
MMSP | 5 |
| 2017 | A generic neural acoustic beamforming architecture for robust multi-channel speech processing
Jahn Heymann, Lukas Drude, Reinhold Häb-Umbach |
Comput. Speech Lang. | 3 |
| 2016 | Blind speech separation based on complex spherical k-mode clusteringabstractWe present an algorithm for clustering complex-valued unit length vectors on the unit hypersphere, which we call complex spherical k-mode clustering, as it can be viewed as a generalization of the spherical k-means algorithm to normalized complex-valued vectors. We show how the proposed algorithm can be derived from the Expectation Maximization algorithm for complex Watson mixture models and prove its applicability in a blind speech separation (BSS) task with real-world room impulse response measurements. It turns out that the proposed spherical k-mode algorithm is on par with other state-of-the-art BSS algorithms in terms of signal-to-inference ratio gains although being far easier to implement and using fewer calculations. Lukas Drude, Christoph Böddeker, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2016 | Neural network based spectral mask estimation for acoustic beamformingabstractWe present a neural network based approach to acoustic beamforming. The network is used to estimate spectral masks from which the Cross-Power Spectral Density matrices of speech and noise are estimated, which in turn are used to compute the beamformer coefficients. The network training is independent of the number and the geometric configuration of the microphones. We further show that it is possible to train the network on clean speech only, avoiding the need for stereo data with separated speech and noise. Two types of networks are evaluated. One small feed-forward network with only one hidden layer and one more elaborated bi-directional Long Short-Term Memory network. We compare our system with different parametric approaches to mask estimation and using different beamforming algorithms. We show that our system yields superior results, both in terms of perceptual speech quality and with respect to speech recognition error rate. The results for the simple feed-forward network are especially encouraging considering its low computational requirements. Jahn Heymann, Lukas Drude, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2016 | A priori SNR Estimation Using a Generalized Decision Directed ApproachabstractIn this contribution we investigate a priori signal-to-noise ratio (SNR) estimation, a crucial component of a single-channel speech enhancement system based on spectral subtraction. The majority of the state-of-the art a priori SNR estimators work in the power spectral domain, which is, however, not confirmed to be the optimal domain for the estimation. Motivated by the generalized spectral subtraction rule, we show how the estimation of the a priori SNR can be formulated in the so called generalized SNR domain. This formulation allows to generalize the widely used decision directed (DD) approach. An experimental investigation with different noise types reveals the superiority of the generalized DD approach over the conventional DD approach in terms of both the mean opinion score - listening quality objective measure and the output global SNR in the medium to high input SNR regime, while we show that the power spectrum is the optimal domain for low SNR. We further develop a parameterization which adjusts the domain of estimation automatically according to the estimated input global SNR. Index Terms: single-channel speech enhancement, a priori SNR estimation, generalized spectral subtraction Aleksej Chinaev, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2016 | On the Appropriateness of Complex-Valued Neural Networks for Speech EnhancementabstractAlthough complex-valued neural networks (CVNNs) â?? networks which can operate with complex arithmetic â?? have been around for a while, they have not been given reconsideration since the breakthrough of deep network architectures. This paper presents a critical assessment whether the novel tool set of deep neural networks (DNNs) should be extended to complex-valued arithmetic. Indeed, with DNNs making inroads in speech enhancement tasks, the use of complex-valued input data, specifically the short-time Fourier transform coefficients, is an obvious consideration. In particular when it comes to performing tasks that heavily rely on phase information, such as acoustic beamforming, complex-valued algorithms are omnipresent. In this contribution we recapitulate backpropagation in CVNNs, develop complex-valued network elements, such as the split-rectified non-linearity, and compare real- and complex-valued networks on a beamforming task. We find that CVNNs hardly provide a performance gain and conclude that the effort of developing the complex-valued counterparts of the building blocks of modern deep or recurrent neural networks can hardly be justified. Lukas Drude, Bhiksha Raj, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2015 | BLSTM supported GEV beamformer front-end for the 3RD CHiME challengeabstractWe present a new beamformer front-end for Automatic Speech Recognition and apply it to the 3rd-CHiME Speech Separation and Recognition Challenge. Without any further modification of the back-end, we achieve a 53% relative reduction of the word error rate over the best baseline enhancement system for the relevant test data set. Our approach leverages the power of a bi-directional Long Short-Term Memory network to robustly estimate soft masks for a subsequent beamforming step. The utilized Generalized Eigenvalue beamforming operation with an optional Blind Analytic Normalization does not rely on a Direction-of-Arrival estimate and can cope with multi-path sound propagation, while at the same time only introducing very limited speech distortions. Our quite simple setup exploits the possibilities provided by simulated training data while still being able to generalize well to the fairly different real data. Finally, combining our front-end with data augmentation and another language model nearly yields a 64 % reduction of the word error rate on the real data test set. Jahn Heymann, Lukas Drude, Aleksej Chinaev, Reinhold Häb-Umbach |
ASRU | 4 |
| 2015 | Unsupervised adaptation of a denoising autoencoder by Bayesian Feature Enhancement for reverberant asr under mismatch conditionsabstractThe parametric Bayesian Feature Enhancement (BFE) and a datadriven Denoising Autoencoder (DA) both bring performance gains in severe single-channel speech recognition conditions. The first can be adjusted to different conditions by an appropriate parameter setting, while the latter needs to be trained on conditions similar to the ones expected at decoding time, making it vulnerable to a mismatch between training and test conditions. We use a DNN backend and study reverberant ASR under three types of mismatch conditions: different room reverberation times, different speaker to microphone distances and the difference between artificially reverberated data and the recordings in a reverberant environment. We show that for these mismatch conditions BFE can provide the targets for a DA. This unsupervised adaptation provides a performance gain over the direct use of BFE and even enables to compensate for the mismatch of real and simulated reverberant data. Jahn Heymann, Reinhold Häb-Umbach, Pavel Golik, Ralf Schlüter |
ICASSP | 2 |
| 2015 | Aligning training modelswith smartphone properties in WiFi fingerprinting based indoor localizationabstractWe are concerned with the so-called fingerprinting method for WiFi-based indoor positioning, where the measured received signal strength index (RSSI) is compared with training data to come up with an estimate of the user's location. We introduce a method for adapting the trained models to the statistics of the RSSI values of the target (testing) WiFi device, which is derived from the Maximum Likelihood Linear Regression (MLLR) framework. By introducing regression classes the assumption of a linear relationship between the RSSI readings of the testing device and the training data is relaxed, leading to superior adaptation performance. Parameter adaptation formulas are derived for the general case of censored and dropped data. While censoring occurs due to the limited sensitivity of WiFi chips, dropping is probably caused by limitations of the operating system of the portable devices. Experiments both on simulated and real-world data demonstrate the effectiveness of the proposed algorithms. Manh Kha Hoang, Joerg Schmalenstroeer, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2015 | Source counting in speech mixtures by nonparametric Bayesian estimation of an infinite Gaussian mixture modelabstractIn this paper we present a source counting algorithm to determine the number of speakers in a speech mixture. In our proposed method, we model the histogram of estimated directions of arrival with a non-parametric Bayesian infinite Gaussian mixture model. As an alternative to classical model selection criteria and to avoid specifying the maximum number of mixture components in advance, a Dirichlet process prior is employed over the mixture components. This allows to automatically determine the optimal number of mixture components that most probably model the observations. We demonstrate by experiments that this model outperforms a parametric approach using a finite Gaussian mixture model with a Dirichlet distribution prior over the mixture weights. Oliver Walter, Lukas Drude, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2015 | On optimal smoothing in minimum statistics based noise trackingabstractNoise tracking is an important component of speech enhance-ment algorithms. Of the many noise trackers proposed, Min-imum Statistics (MS) is a particularly popular one due to its simple parameterization and at the same time excellent perfor-mance. In this paper we propose to further reduce the number of MS parameters by giving an alternative derivation of an optimal smoothing constant. At the same time the noise tracking perfor-mance is improved as is demonstrated by experiments employ-ing speech degraded by various noise types and at different SNR values. Index Terms: speech enhancement, noise tracking, optimal smoothing Aleksej Chinaev, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2015 | Semantic analysis of spoken input using Markov logic networksabstractWe present a semantic analysis technique for spoken input us-ing Markov Logic Networks (MLNs). MLNs combine graphi-cal models with first-order logic. They are particularly suitable for providing inference in the presence of inconsistent and in-complete data, which are typical of an automatic speech rec-ognizer’s (ASR) output in the presence of degraded speech. The target application is a speech interface to a home automa-tion system to be operated by people with speech impairments, where the ASR output is particularly noisy. In order to cater for dysarthric speech with non-canonical phoneme realizations, acoustic representations of the input speech are learned in an unsupervised fashion. While training data transcripts are not required for the acoustic model training, the MLN training re-quires supervision, however, at a rather loose and abstract level. Results on two databases, one of them for dysarthric speech, show that MLN-based semantic analysis clearly outperforms baseline approaches employing non-negative matrix factoriza-tion, multinomial naive Bayes models, or support vector ma-chines. Vladimir Despotovic, Oliver Walter, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2015 | Typicality and emotion in the voice of children with autism spectrum condition: evidence across three languagesabstractOnly a few studies exist on automatic emotion analysis of speech from children with Autism Spectrum Conditions (ASC). Out of these, some preliminary studies have recently focused on comparing the relevance of selected prosodic features against large sets of acoustic, spectral, and cepstral features; however, no study so far provided a comparison of performances across dif-ferent languages. The present contribution aims to fill this white spot in the literature and provide insight by extensive evaluations carried out on three databases of prompted phrases collected in English, Swedish, and Hebrew, inducing nine emotion categories embedded in short-stories. The datasets contain speech of chil-dren with ASC and typically developing children under the same conditions. We evaluate automatic diagnosis and recognition of emotions in atypical childrens voice over the nine categories including binary valence/arousal discrimination. Erik Marchi, Björn W. Schuller, Simon Baron-Cohen, Ofer Golan, Sven Bölte, Prerna Arora, Reinhold Häb-Umbach |
INTERSPEECH | 7 |
| 2015 | A combined hardware-software approach for acoustic sensor network synchronization
Joerg Schmalenstroeer, Patrick Jebramcik, Reinhold Häb-Umbach |
Signal Process. | 3 |
| 2014 | Source counting in speech mixtures using a variational EM approach for complex WATSON mixture modelsabstractIn this contribution we derive a variational EM (VEM) algorithm for model selection in complex Watson mixture models, which have been recently proposed as a model of the distribution of normalized microphone array signals in the short-time Fourier transform domain. The VEM algorithm is applied to count the number of active sources in a speech mixture by iteratively estimating the mode vectors of the Watson distributions and suppressing the signals from the corresponding directions. A key theoretical contribution is the derivation of the MMSE estimate of a quadratic form involving the mode vector of the Watson distribution. The experimental results demonstrate the effectiveness of the source counting approach at moderately low SNR. It is further shown that the VEM algorithm is more robust with respect to used threshold values. Lukas Drude, Aleksej Chinaev, Dang Hai Tran Vu, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2014 | Iterative Bayesian word segmentation for unsupervised vocabulary discovery from phoneme latticesabstractIn this paper we present an algorithm for the unsupervised segmentation of a lattice produced by a phoneme recognizer into words. Using a lattice rather than a single phoneme string accounts for the uncertainty of the recognizer about the true label sequence. An example application is the discovery of lexical units from the output of an error-prone phoneme recognizer in a zero-resource setting, where neither the lexicon nor the language model (LM) is known. We propose a computationally efficient iterative approach, which alternates between the following two steps: First, the most probable string is extracted from the lattice using a phoneme LM learned on the segmentation result of the previous iteration. Second, word segmentation is performed on the extracted string using a word and phoneme LM which is learned alongside the new segmentation. We present results on lattices produced by a phoneme recognizer on the WSJ-CAM0 dataset. We show that our approach delivers superior segmentation performance than an earlier approach found in the literature, in particular for higher-order language models. Jahn Heymann, Oliver Walter, Reinhold Häb-Umbach, Bhiksha Raj |
ICASSP | 3 |
| 2014 | A gossiping approach to sampling clock synchronization in wireless acoustic sensor networksabstractIn this paper we present an approach for synchronizing the sampling clocks of distributed microphones over a wireless network. The proposed system uses a two stage procedure. It first employs a two-way message exchange algorithm to estimate the clock phase and frequency difference between two nodes and then uses a gossiping algorithm to estimate a virtual master clock, to which all sensor nodes synchronize. Simulation results are presented for networks of different topology and size, showing the effectiveness of our approach. Joerg Schmalenstroeer, Patrick Jebramcik, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2014 | An evaluation of unsupervised acoustic model training for a dysarthric speech interfaceabstractCopyright © 2014 ISCA. In this paper, we investigate unsupervised acoustic model training approaches for dysarthric-speech recognition. These models are first, frame-based Gaussian posteriorgrams, obtained from Vector Quantization (VQ), second, so-called Acoustic Unit Descriptors (AUDs), which are hidden Markov models of phone-like units, that are trained in an unsupervised fashion, and, third, posteriorgrams computed on the AUDs. Experiments were carried out on a database collected from a home automation task and containing nine speakers, of which seven are considered to utter dysarthric speech. All unsupervised modeling approaches delivered significantly better recognition rates than a speaker-independent phoneme recognition baseline, showing the suitability of unsupervised acoustic model training for dysarthric speech. While the AUD models led to the most compact representation of an utterance for the subsequent semantic inference stage, posteriorgram-based representations resulted in higher recognition rates, with the Gaussian posteriorgram achieving the highest slot filling F-score of 97.02%. Oliver Walter, Vladimir Despotovic, Reinhold Häb-Umbach, Jort F. Gemmeke, Bart Ons, Hugo Van hamme |
INTERSPEECH | 3 |
| 2014 | A New Observation Model in the Logarithmic Mel Power Spectral Domain for the Automatic Recognition of Noisy Reverberant SpeechabstractIn this contribution we present a theoretical and experimental investigation into the effects of reverberation and noise on features in the logarithmic mel power spectral domain, an intermediate stage in the computation of the mel frequency cepstral coefficients, prevalent in automatic speech recognition (ASR). Gaining insight into the complex interaction between clean speech, noise, and noisy reverberant speech features is essential for any ASR system to be robust against noise and reverberation present in distant microphone input signals. The findings are gathered in a probabilistic formulation of an observation model which may be used in model-based feature compensation schemes. The proposed observation model extends previous models in three major directions: First, the contribution of additive background noise to the observation error is explicitly taken into account. Second, an energy compensation constant is introduced which ensures an unbiased estimate of the reverberant speech features, and, third, a recursive variant of the observation model is developed resulting in reduced computational complexity when used in model-based feature compensation. The experimental section is used to evaluate the accuracy of the model and to describe how its parameters can be determined from test data. Volker Leutnant, Alexander Krueger, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | An Overview of Noise-Robust Automatic Speech RecognitionabstractNew waves of consumer-centric applications, such as voice search and voice interaction with mobile devices and home entertainment systems, increasingly require automatic speech recognition (ASR) to be robust to the full range of real-world noise and other acoustic distorting conditions. Despite its practical importance, however, the inherent links between and distinctions among the myriad of methods for noise-robust ASR have yet to be carefully studied in order to advance the field further. To this end, it is critical to establish a solid, consistent, and common mathematical foundation for noise-robust ASR, which is lacking at present. This article is intended to fill this gap and to provide a thorough overview of modern noise-robust techniques for ASR developed over the past 30 years. We emphasize methods that are proven to be successful and that are likely to sustain or expand their future applicability. We distill key insights from our comprehensive overview in this field and take a fresh look at a few old problems, which nevertheless are still highly relevant today. Specifically, we have analyzed and categorized a wide range of noise-robust techniques using five different criteria: 1) feature-domain vs. model-domain processing, 2) the use of prior knowledge about the acoustic environment distortion, 3) the use of explicit environment-distortion models, 4) deterministic vs. uncertainty processing, and 5) the use of acoustic models trained jointly with the same feature enhancement or model adaptation process used in the testing stage. With this taxonomy-oriented review, we equip the reader with the insight to choose among techniques and with the awareness of the performance-complexity tradeoffs. The pros and cons of using different noise-robust ASR techniques in practical application scenarios are provided as a guide to interested practitioners. The current challenges and future research directions in this field is also carefully analyzed. Jinyu Li 0001, Li Deng 0001, Yifan Gong 0001, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Unsupervised word segmentation from noisy inputabstractIn this paper we present an algorithm for the unsupervised segmentation of a character or phoneme lattice into words. Using a lattice at the input rather than a single string accounts for the uncertainty of the character/phoneme recognizer about the true label sequence. An example application is the discovery of lexical units from the output of an error-prone phoneme recognizer in a zero-resource setting, where neither the lexicon nor the language model is known. Recently a Weighted Finite State Transducer (WFST) based approach has been published which we show to suffer from an issue: language model probabilities of known words are computed incorrectly. Fixing this issue leads to greatly improved precision and recall rates, however at the cost of increased computational complexity. It is therefore practical only for single input strings. To allow for a lattice input and thus for errors in the character/phoneme recognizer, we propose a computationally efficient suboptimal two-stage approach, which is shown to significantly improve the word segmentation performance compared to the earlier WFST approach. Jahn Heymann, Oliver Walter, Reinhold Häb-Umbach, Bhiksha Raj |
ASRU | 3 |
| 2013 | A hierarchical system for word discovery exploiting DTW-based initializationabstractDiscovering the linguistic structure of a language solely from spoken input asks for two steps: phonetic and lexical discovery. The first is concerned with identifying the categorical subword unit inventory and relating it to the underlying acoustics, while the second aims at discovering words as repeated patterns of subword units. The hierarchical approach presented here accounts for classification errors in the first stage by modelling the pronunciation of a word in terms of subword units probabilistically: a hidden Markov model with discrete emission probabilities, emitting the observed subword unit sequences. We describe how the system can be learned in a completely unsupervised fashion from spoken input. To improve the initialization of the training of the word pronunciations, the output of a dynamic time warping based acoustic pattern discovery system is used, as it is able to discover similar temporal sequences in the input data. This improved initialization, using only weak supervision, has led to a 40% reduction in word error rate on a digit recognition task. Oliver Walter, Timo Korthals, Reinhold Häb-Umbach, Bhiksha Raj |
ASRU | 3 |
| 2013 | GMM-based significance decodingabstractThe accuracy of automatic speech recognition systems in noisy and reverberant environments can be improved notably by exploiting the uncertainty of the estimated speech features using so-called uncertainty-of-observation techniques. In this paper, we introduce a new Bayesian decision rule that can serve as a mathematical framework from which both known and new uncertainty-of-observation techniques can be either derived or approximated. The new decision rule in its direct form leads to the new significance decoding approach for Gaussian mixture models, which results in better performance compared to standard uncertainty-of-observation techniques in different additive and convolutive noise scenarios. Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa, Volker Leutnant, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2013 | Map-based estimation of the parameters of a Gaussian Mixture Model in the presence of noisy observationsabstractIn this contribution we derive the Maximum A-Posteriori (MAP) estimates of the parameters of a Gaussian Mixture Model (GMM) in the presence of noisy observations. We assume the distortion to be white Gaussian noise of known mean and variance. An approximate conjugate prior of the GMM parameters is derived allowing for a computationally efficient implementation in a sequential estimation framework. Simulations on artificially generated data demonstrate the superiority of the proposed method compared to the Maximum Likelihood technique and to the ordinary MAP approach, whose estimates are corrected by the known statistics of the distortion in a straightforward manner. Aleksej Chinaev, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2013 | Improved single-channel nonstationary noise tracking by an optimized MAP-based postprocessorabstractIn this paper we present an improved version of the recently proposed Maximum A-Posteriori (MAP) based noise power spectral density estimator. An empirical bias compensation and bandwidth adjustment reduce bias and variance of the noise variance estimates. The main advantage of the MAP-based postprocessor is its low estimation variance. The estimator is employed in the second stage of a two-stage single-channel speech enhancement system, where eight different state-of-the-art noise tracking algorithms were tested in the first stage. While the postprocessor hardly affects the results in stationary noise scenarios, it becomes the more effective the more nonstationary the noise is. The proposed postprocessor was able to improve all systems in babble noise w.r.t. the perceptual evaluation of speech quality performance. Aleksej Chinaev, Reinhold Häb-Umbach, Jalal Taghia, Rainer Martin 0001 |
ICASSP | 2 |
| 2013 | Parameter estimation and classification of censored Gaussian data with application to WiFi indoor positioningabstractIn this paper, we consider the Maximum Likelihood (ML) estimation of the parameters of a GAUSSIAN in the presence of censored, i.e., clipped data. We show that the resulting Expectation Maximization (EM) algorithm delivers virtually biasfree and efficient estimates, and we discuss its convergence properties. We also discuss optimal classification in the presence of censored data. Censored data are frequently encountered in wireless LAN positioning systems based on the fingerprinting method employing signal strength measurements, due to the limited sensitivity of the portable devices. Experiments both on simulated and real-world data demonstrate the effectiveness of the proposed algorithms. Manh Kha Hoang, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2013 | DOA-based microphone array postion self-calibration using circular statisticsabstractIn this paper we propose an approach to retrieve the absolute geometry of an acoustic sensor network, consisting of spatially distributed microphone arrays, from reverberant speech input. The calibration relies on direction of arrival measurements of the individual arrays. The proposed calibration algorithm is derived from a maximum-likelihood approach employing circular statistics. Since a sensor node consists of a microphone array with known intra-array geometry, we are able to obtain an absolute geometry estimate, including angles and distances. Simulation results demonstrate the effectiveness of the approach. Florian Jacob, Joerg Schmalenstroeer, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2013 | Using the turbo principle for exploiting temporal and spectral correlations in speech presence probability estimationabstractIn this paper we present a speech presence probability (SPP) estimation algorithmwhich exploits both temporal and spectral correlations of speech. To this end, the SPP estimation is formulated as the posterior probability estimation of the states of a two-dimensional (2D) Hidden Markov Model (HMM). We derive an iterative algorithm to decode the 2D-HMM which is based on the turbo principle. The experimental results show that indeed the SPP estimates improve from iteration to iteration, and further clearly outperform another state-of-the-art SPP estimation algorithm. Dang Hai Tran Vu, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2013 | Bayesian Feature Enhancement for Reverberation and Noise Robust Speech RecognitionabstractIn this contribution we extend a previously proposed Bayesian approach for the enhancement of reverberant logarithmic mel power spectral coefficients for robust automatic speech recognition to the additional compensation of background noise. A recently proposed observation model is employed whose time-variant observation error statistics are obtained as a side product of the inference of the a posteriori probability density function of the clean speech feature vectors. Further a reduction of the computational effort and the memory requirements are achieved by using a recursive formulation of the observation model. The performance of the proposed algorithms is first experimentally studied on a connected digits recognition task with artificially created noisy reverberant data. It is shown that the use of the time-variant observation error model leads to a significant error rate reduction at low signal-to-noise ratios compared to a time-invariant model. Further experiments were conducted on a 5000 word task recorded in a reverberant and noisy environment. A significant word error rate reduction was obtained demonstrating the effectiveness of the approach on real-world data. Volker Leutnant, Alexander Krueger, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Improved noise power spectral density tracking by a MAP-based postprocessorabstractIn this paper we present a novel noise power spectral density tracking algorithm and its use in single-channel speech enhancement. It has the unique feature that it is able to track the noise statistics even if speech is dominant in a given time-frequency bin. As a consequence it can follow non-stationary noise superposed by speech, even in the critical case of rising noise power. The algorithm requires an initial estimate of the power spectrum of speech and is thus meant to be used as a postprocessor to a first speech enhancement stage. An experimental comparison with a state-of-the-art noise tracking algorithm demonstrates lower estimation errors under low SNR conditions and smaller fluctuations of the estimated values, resulting in improved speech quality as measured by PESQ scores. Aleksej Chinaev, Alexander Krueger, Dang Hai Tran Vu, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2012 | Bayesian Feature Enhancement for ASR of Noisy Reverberant Real-World DataabstractIn this contribution we investigate the effectiveness of Bayesian feature enhancement (BFE) on a medium-sized recognition task containing real-world recordings of noisy reverberant speech. BFE employs a very coarse model of the acoustic impulse response (AIR) from the source to the microphone, which has been shown to be effective if the speech to be recognized has been generated by artificially convolving nonreverberant speech with a constant AIR. Here we demonstrate that the model is also appropriate to be used in feature enhancement of true recordings of noisy reverberant speech. On the Multi-Channel Wall Street Journal Audio Visual corpus (MC-WSJ-AV) the word error rate is cut in half to 41.9 percent compared to the ETSI Standard Front-End using as input the signal of a single distant microphone with a single recognition pass. Alexander Krueger, Oliver Walter, Volker Leutnant, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2011 | MAP-based estimation of the parameters of non-stationary Gaussian processes from noisy observationsabstractThe paper proposes a modification of the standard maximum a posteriori (MAP) method for the estimation of the parameters of a Gaussian process for cases where the process is superposed by additive Gaussian observation errors of known variance. Simulations on artificially generated data demonstrate the superiority of the proposed method. While reducing to the ordinary MAP approach in the absence of observation noise, the improvement becomes the more pronounced the larger the variance of the observation noise. The method is further extended to track the parameters in case of non-stationary Gaussian processes. Alexander Krueger, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2011 | A Versatile Gaussian Splitting Approach to Non-Linear State Estimation and its Application to Noise-Robust ASRabstractIn this work, a splitting and weighting scheme that allows for splitting a Gaussian density into a Gaussian mixture density (GMM) is extended to allow the mixture components to be arranged along arbitrary directions. The parameters of the Gaussian mixture are chosen such that the GMM and the original Gaussian still exhibit equal central moments up to an order of four. The resulting mixtures{\rq} covariances will have eigenvalues that are smaller than those of the covariance of the original distribution, which is a desirable property in the context of non-linear state estimation, since the underlying assumptions of the extended K ALMAN filter are better justified in this case. Application to speech feature enhancement in the context of noise-robust automatic speech recognition reveals the beneficial properties of the proposed approach in terms of a reduced word error rate on the Aurora 2 recognition task. Volker Leutnant, Alexander Krueger, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2011 | Unsupervised Learning of Acoustic Events Using Dynamic Time Warping and Hierarchical K-Means++ ClusteringabstractIn this paper we propose to jointly consider Segmental Dynamic Time Warping and distance clustering for the unsupervised learning of acoustic events. As a result, the computational complexity increases only linearly with the dababase size compared to a quadratic increase in a sequential setup, where all pairwise SDTW distances between segments are computed prior to clustering. Further, we discuss options for seed value selection for clustering and show that drawing seeds with a probability proportional to the distance from the already drawn seeds, known as K-means++ clustering, results in a significantly higher probability of finding representatives of each of the underlying classes, compared to the commonly used draws from a uniform distribution. Experiments are performed on an acoustic event classification and an isolated digit recognition task, where on the latter the final word accuracy approaches that of supervised training. Index Terms: unsupervised, clustering, acoustic events 1. Joerg Schmalenstroeer, Markus Bartek, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2011 | Unsupervised Geometry Calibration of Acoustic Sensor Networks Using Source CorrespondencesabstractIn this paper we propose a procedure for estimating the geometric configuration of an arbitrary acoustic sensor placement. It determines the position and the orientation of microphone arrays in 2D while locating a source by direction-of-arrival (DoA) estimation. Neither artificial calibration signals nor unnatural user activity are required. The problem of scale indeterminacy inherent to DoA-only observations is solved by adding time difference of arrival (TDOA) measurements. The geometry calibration method is numerically stable and delivers precise results in moderately reverberated rooms. Simulation results are confirmed by laboratory experiments. Joerg Schmalenstroeer, Florian Jacob, Reinhold Häb-Umbach, Marius H. Hennecke, Gernot A. Fink |
INTERSPEECH | 3 |
| 2011 | On Initial Seed Selection for Frequency Domain Blind Speech SeparationabstractIn this paper we address the problem of initial seed selection for frequency domain iterative blind speech separation (BSS) algorithms. The derivation of the seeding algorithm is guided by the goal to select samples which are likely to be caused by source activity and not by noise and at the same time originate from different sources. The proposed algorithm has moderate computational complexity and finds better seed values than alternative schemes, as is demonstrated by experiments on the database of the SiSEC2010 challenge. Dang Hai Tran Vu, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2011 | Speech Enhancement With a GSC-Like Structure Employing Eigenvector-Based Transfer Function Ratios EstimationabstractIn this paper, we present a novel blocking matrix and fixed beamformer design for a generalized sidelobe canceler for speech enhancement in a reverberant enclosure. They are based on a new method for estimating the acoustical transfer function ratios in the presence of stationary noise. The estimation method relies on solving a generalized eigenvalue problem in each frequency bin. An adaptive eigenvector tracking utilizing the power iteration method is employed and shown to achieve a high convergence speed. Simulation results demonstrate that the proposed beamformer leads to better noise and interference reduction and reduced speech distortions compared to other blocking matrix designs from the literature. Alexander Krueger, Ernst Warsitz, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Blind speech separation employing directional statistics in an Expectation Maximization frameworkabstractIn this paper we propose to employ directional statistics in a complex vector space to approach the problem of blind speech separation in the presence of spatially correlated noise. We interpret the values of the short time Fourier transform of the microphone signals to be draws from a mixture of complexWatson distributions, a probabilistic model which naturally accounts for spatial aliasing. The parameters of the density are related to the a priori source probabilities, the power of the sources and the transfer function ratios from sources to sensors. Estimation formulas are derived for these parameters by employing the Expectation Maximization (EM) algorithm. The E-step corresponds to the estimation of the source presence probabilities for each time-frequency bin, while the M-step leads to a maximum signal-to-noise ratio (MaxSNR) beamformer in the presence of uncertainty about the source activity. Experimental results are reported for an implementation in a generalized sidelobe canceller (GSC) like spatial beamforming configuration for 3 speech sources with significant coherent noise in reverberant environments, demonstrating the usefulness of the novel modeling framework. Dang Hai Tran Vu, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2010 | On the exploitation of hidden Markov models and linear dynamic models in a hybrid decoder architecture for continuous speech recognitionabstractLinear dynamic models (LDMs) have been shown to be a viable alternative to hidden Markov models (HMMs) on small-vocabulary recognition tasks, such as phone classification. In this paper we investigate various statistical model combination approaches for a hybrid HMM-LDM recognizer, resulting in a phone classification performance that outperforms the best individual classifier. Further, we report on continuous speech recognition experiments on the AURORA4 corpus, where the model combination is carried out on wordgraph rescoring. While the hybrid system improves the HMM system in the case of monophone HMMs, the performance of the triphone HMM model could not be improved by monophone LDMs, asking for the need to introduce context-dependency also in the LDM model inventory. Volker Leutnant, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2010 | Ungrounded independent non-negative factor analysisabstractWe describe an algorithm that performs regularized non-negative matrix factorization (NMF) to find independent components in nonnegative data. Previous techniques proposed for this purpose require the data to be grounded, with support that goes down to 0 along each dimension. In our work, this requirement is eliminated. Based on it, we present a technique to find a low-dimensional decomposition of spectrograms by casting it as a problem of discovering independent non-negative components from it. The algorithm itself is implemented as regularized non-negative matrix factorization (NMF). Unlike other ICA algorithms, this algorithm computes the mixing matrix rather than an unmixing matrix. This algorithm provides a better decomposition than standard NMF when the underlying sources are independent. It makes better use of additional observation streams than previous nonnegative ICA algorithms. Index Terms — matrix decomposition, ICA 1. Bhiksha Raj, Kevin W. Wilson, Alexander Krueger, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2010 | Model-Based Feature Enhancement for Reverberant Speech RecognitionabstractIn this paper, we present a new technique for automatic speech recognition (ASR) in reverberant environments. Our approach is aimed at the enhancement of the logarithmic Mel power spectrum, which is computed at an intermediate stage to obtain the widely used Mel frequency cepstral coefficients (MFCCs). Given the reverberant logarithmic Mel power spectral coefficients (LMPSCs), a minimum mean square error estimate of the clean LMPSCs is computed by carrying out Bayesian inference. We employ switching linear dynamical models as ana priorimodel for the dynamics of the clean LMPSCs. Further, we derive a stochastic observation model which relates the clean to the reverberant LMPSCs through a simplified model of the room impulse response (RIR). This model requires only two parameters, namely RIR energy and reverberation time, which can be estimated from the captured microphone signal. The performance of the proposed enhancement technique is studied on the AURORA5 database and compared to that of constrained maximum-likelihood linear regression (CMLLR). It is shown by experimental results that our approach significantly outperforms CMLLR and that up to 80% of the errors caused by the reverberation are recovered. In addition to the fact that the approach is compatible with the standard MFCC feature vectors, it leaves the ASR back-end unchanged. It is of moderate computational complexity and suitable for real time applications. Alexander Krueger, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Model based feature enhancement for automatic speech recognition in reverberant environmentsabstractIn this paper we present a new feature space dereverberation technique for automatic speech recognition. We derive an expression for the dependence of the reverberant speech features in the log-mel spectral domain on the non-reverberant speech features and the room impulse response. The obtained observation model is used for a model based speech enhancement based on Kalman filtering. The performance of the proposed enhancement technique is studied on the AURORA5 database. In our currently best configuration, which includes uncertainty decoding, the number of recognition errors is approximately halved compared to the recognition of unprocessed speech. Alexander Krueger, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2009 | An analytic derivation of a phase-sensitive observation model for noise robust speech recognitionabstractIn this paper we present an analytic derivation of the moments of the phase factor between clean speech and noise cepstral or log-mel-spectral feature vectors. The development shows, among others, that the probability density of the phase factor is of sub-Gaussian nature and that it is independent of the noise type and the signal-to-noise ratio, however dependent on the mel filter bank index. Further we show how to compute the contribution of the phase factor to both the mean and the variance of the noisy speech observation likelihood, which relates the speech and noise feature vectors to those of noisy speech. The resulting phase-sensitive observation model is then used in model-based speech feature enhancement, leading to significant improvements in word accuracy on the AURORA2 database. Index Terms: model-based feature enhancement, phasesensitive observation model, phase factor distribution Volker Leutnant, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2009 | Fusing audio and video information for online speaker diarizationabstractIn this paper we present a system for identifying and localizing speakers using distant microphone arrays and a steerable pantilt-zoom camera. The scenario at hand assumes audio streams to be processed in real-time to get the diarization information “who spokes when and where ” with only short delays. Our new idea is to fuse the acoustical and visual observations directly within the Viterbi decoder to improve the diarization process. In contrast to standard Viterbi decoder implementations, we use a time variant transition matrix generated from speaker change hypotheses and location information. This allows a simultaneous segmentation and classification of the audio stream. Experiments show, that video information enables a substantial improvement of the diarization results. Index Terms: speaker diarization, face identification, acoustic scene analysis Joerg Schmalenstroeer, Martin Kelling, Volker Leutnant, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2009 | Joint Parameter Estimation and Tracking in a Multi-Stage Kalman Filter for Vehicle PositioningabstractIn this paper we present a novel vehicle tracking method which is based on multi-stage Kalman filtering of GPS and IMU sensor data. After individual Kalman filtering of GPS and IMU measurements the estimates of the orientation of the vehicle are combined in an optimal manner to improve the robustness towards drift errors. The tracking algorithm incorporates the estimation of time-variant covariance parameters by using an iterative block Expectation-Maximization algorithm to account for time-variant driving conditions and measurement quality. The proposed system is compared to an interacting multiple model approach (IMM) and achieves improved localization accuracy at lower computational complexity. Furthermore we show how the joint parameter estimation and localizaiton can be conducted with streaming input data to be able to track vehicles in a real driving environment. Maik Bevermeier, Sven Peschke, Reinhold Häb-Umbach |
VTC Spring | 3 |
| 2009 | Approaches to Iterative Speech Feature Enhancement and RecognitionabstractIn automatic speech recognition, hidden Markov models (HMMs) are commonly used for speech decoding, while switching linear dynamic models (SLDMs) can be employed for a preceding model-based speech feature enhancement. In this paper, these model types are combined in order to obtain a novel iterative speech feature enhancement and recognition architecture. It is shown that speech feature enhancement with SLDMs can be improved by feeding back information from the HMM to the enhancement stage. Two different feedback structures are derived. In the first, the posteriors of the HMM states are used to control the model probabilities of the SLDMs, while in the second they are employed to directly influence the estimate of the speech feature distribution. Both approaches lead to improvements in recognition accuracy both on the AURORA2 and AURORA4 databases compared to non-iterative speech feature enhancement with SLDMs. It is also shown that a combination with uncertainty decoding further enhances performance. Stefan Windmann, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Parameter Estimation of a State-Space Model of Noise for Robust Speech RecognitionabstractIn this paper, parameter estimation of a state-space model of noise or noisy speech cepstra is investigated. A blockwise EM algorithm is derived for the estimation of the state and observation noise covariance from noise-only input data. It is supposed to be used during the offline training mode of a speech recognizer. Further a sequential online EM algorithm is developed to adapt the observation noise covariance on noisy speech cepstra at its input. The estimated parameters are then used in model-based speech feature enhancement for noise-robust automatic speech recognition. Experiments on the AURORA4 database lead to improved recognition results with a linear state model compared to the assumption of stationary noise. Stefan Windmann, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Speech enhancement with a new generalized eigenvector blocking matrix for application in a generalized sidelobe cancellerabstractThe generalized sidelobe canceller by Griffith and Jim is a robust beamforming method to enhance a desired (speech) signal in the presence of stationary noise. Its performance depends to a high degree on the construction of the blocking matrix which produces noise reference signals for the subsequent adaptive interference canceller. Especially in reverberated environments the beamformer may suffer from signal leakage and reduced noise suppression. In this paper a new blocking matrix is proposed. It is based on a generalized eigenvalue problem whose solution provides an indirect estimation of the transfer functions from the source to the sensors. The quality of the new generalized eigenvector blocking matrix is studied in simulated rooms with different reverberation times and is compared to alternatives proposed in the literature. Ernst Warsitz, Alexander Krueger, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2008 | Modeling the dynamics of speech and noise for speech feature enhancement in ASRabstractIn this paper a switching linear dynamical model (SLDM) approach for speech feature enhancement is improved by employing more accurate models for the dynamics of speech and noise. The model of the clean speech feature trajectory is improved by augmenting the state vector to capture information derived from the delta features. Further a hidden noise state variable is introduced to obtain a more elaborated model for the noise dynamics. Approximate Bayesian inference in the SLDM is carried out by a bank of extended Kalman filters, whose outputs are combined according to the a posteriori probability of the individual state models. Experimental results on the AURORA2 database show improved recognition accuracy. Stefan Windmann, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2008 | A Novel Uncertainty Decoding Rule With Applications to Transmission Error Robust Speech RecognitionabstractIn this paper, we derive an uncertainty decoding rule for automatic speech recognition (ASR), which accounts for both corrupted observations and inter-frame correlation. The conditional independence assumption, prevalent in hidden Markov model-based ASR, is relaxed to obtain a clean speech posterior that is conditioned on the complete observed feature vector sequence. This is a more informative posterior than one conditioned only on the current observation. The novel decoding is used to obtain a transmission-error robust remote ASR system, where the speech capturing unit is connected to the decoder via an error-prone communication network. We show how the clean speech posterior can be computed for communication links being characterized by either bit errors or packet loss. Recognition results are presented for both distributed and network speech recognition, where in the latter case common voice-over-IP codecs are employed. Valentin Ion, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | OFDM Channel Estimation Based on Combined Estimation in Time and Frequency DomainabstractIn this paper we present a novel channel impulse response estimation technique for block-oriented OFDM transmission based on combining estimators: the estimates provided by a Kalman filter operating in the time domain and a Wiener filter in the frequency domain are optimally combined by taking into account their estimated error covariances. The resulting estimator turns out to be identical to the MAP estimator of correlated jointly Gaussian mean vectors. Different variants of the proposed scheme are experimentally investigated in an EEEE 802.11a-like system setup. They compare favourably with known approaches from the literature resulting in reduced mean square estimation error and bit error rate. Further, robustness and complexity issues are discussed. Reinhold Häb-Umbach, Maik Bevermeier |
ICASSP (3) | 1 |
| 2007 | Multi-resolution soft features for channel-robust distributed speech recognitionabstractIn this paper we introduce soft features of variable resolution for robust distributed speech recognition over channels exhibiting packet losses. The underlying rationale is that lost feature vectors can never be reconstructed perfectly and therefore reconstruction is carried out at a lower resolution than the resolution of the originally sent features. By doing so, enormous reductions in computational effort can be achieved at a graceful or even no degradation in word accuracy. In experiments conducted on the Aurora II database we obtained for example a reduction of a factor of 30 in computation time for the reconstruction of the soft features without an effect on the word error rate. The proposed method is fully compatible with the ETSI DSR standard, as there are no changes involved in the front-end processing and the transmission format. Index Terms: distributed speech recognition, error concealment, uncertainty decoding, soft feature Valentin Ion, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2007 | Joint speaker segmentation, localization and identification for streaming audioabstractIn this paper we investigate the problem of identifying and localizing speakers with distant microphone arrays, thus extending the classical speaker diarization task to answer the question “who spoke when and where”. We consider a streaming audio scenario, where the diarization output is to be generated in realtime with as low latency as possible. Rather than carrying out the individual segmentation and classification tasks (speech detection, change detection, gender/speaker classification) sequentially, we propose a simultaneous segmentation and classification by applying a Viterbi decoder. It uses a transition matrix estimated online from position information and speaker change hypotheses, instead of fixed transition probabilites. This avoids early hard decisions and is shown to outperform the sequential approach. Index Terms: speaker diarization, acoustic scene analysis, Viterbi decoder Joerg Schmalenstroeer, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2007 | Blind adaptive principal eigenvector beamforming for acoustical source separationabstractFor separating multiple speech signals given a convolutive mixture, time-frequency sparseness of the speech sources can be exploited. In this paper we present a multi-channel source separation method based on the concept of approximate disjoint orthogonality of speech signals. Unlike binary masking of singlechannel signals as e.g. applied in the DUET algorithm we use a likelihood mask to control the adaptation of blind principal eigenvector beamformers. Furthermore orthogonal projection of the adapted beamformer filters leads to mutually orthogonal filter coefficients thus enhancing the demixing performance. Experimental results in terms of the achievable signalto-interference ratio (SIR) and a perceptual speech quality measure are given for the proposed method and are compared to the DUET algorithm. Ernst Warsitz, Reinhold Häb-Umbach, Dang Hai Tran Vu |
INTERSPEECH | 2 |
| 2007 | An approach to iterative speech feature enhancement and recognition
Stefan Windmann, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2007 | Blind Acoustic Beamforming Based on Generalized Eigenvalue DecompositionabstractMaximizing the output signal-to-noise ratio (SNR) of a sensor array in the presence of spatially colored noise leads to a generalized eigenvalue problem. While this approach has extensively been employed in narrowband (antenna) array beamforming, it is typically not used for broadband (microphone) array beamforming due to the uncontrolled amount of speech distortion introduced by a narrowband SNR criterion. In this paper, we show how the distortion of the desired signal can be controlled by a single-channel post-filter, resulting in a performance comparable to the generalized minimum variance distortionless response beamformer, where arbitrary transfer functions relate the source and the microphones. Results are given both for directional and diffuse noise. A novel gradient ascent adaptation algorithm is presented, and its good convergence properties are experimentally revealed by comparison with alternatives from the literature. A key feature of the proposed beamformer is that it operates blindly, i.e., it neither requires knowledge about the array geometry nor an explicit estimation of the transfer functions from source to sensors or the direction-of-arrival. Ernst Warsitz, Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | An Inexpensive Packet Loss Compensation Scheme for Distributed Speech Recognition Based on Soft-FeaturesabstractSoft-feature based speech recognition, which is an example of uncertainty decoding, has been proven to be a robust error mitigation method for distributed speech recognition over wireless channels exhibiting bit errors. In this paper we extend this concept to packet-oriented transmissions. The a posteriori probability density function of the lost feature vector, given the closest received neighbours, is computed. In the experiments, the nearest frame repetition, which is shown to be equivalent to the MAP estimate, outperforms the MMSE estimate for long bursts. Taking the variance into account at the speech recognition stage results in superior performance compared to classical schemes using point estimates. A computationally and memory efficient implementation of the proposed packet loss compensation scheme based on table lookup is presented Valentin Ion, Reinhold Häb-Umbach |
ICASSP (1) | 2 |
| 2006 | Iterative Speech Enhancement using a Non-Linear Dynamic State Model of Speech and its ParametersabstractA marginalized particle filter is proposed for performing single channel speech enhancement with a non-linear dynamic state model. The system consists of a particle filter for tracking line spectral pair (LSP) parameters and a Kalman filter per particle for speech enhancement. The state model for the LSPs has been learnt on clean speech training data. In our approach parameters and speech samples are processed at different time scales by assuming the parameters to be constant for small blocks of data. Further enhancement is obtained by an iteration which can be applied on these small blocks. The experiments show that similar SNR gains are obtained as with the Kalman-LM-iterative algorithm. However better values of the noise level and the log-spectral distance are achieved Stefan Windmann, Reinhold Häb-Umbach |
ICASSP (1) | 2 |
| 2006 | Improved source modeling and predictive classification for channel robust speech recognitionabstractThe accuracy of distributed speech recognition has been shown to be very sensitive to errors occurring during transmission. One reason for this is that the classifier, usually trained under error free conditions, is unable to cope with the mismatch between an error free and error prone channel. In this paper we present a novel decision rule for classification which is able to account for channel errors. To achieve this, the classical Bayesian speech recognition approach has been reformulated for the server side, where the observation is known only to the extent, as is given by its a posteriori density function. We present a method to estimate the a posteriori density which is based on a Markov model of the source, which captures correlations of both static and dynamic features. A practical implementation is given, accompanied by experimental results for distributed speech recognition over an IP-network. Index Terms: distributed speech recognition, channel robustness, predictive classification. Valentin Ion, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2006 | Online speaker change detection by combining BIC with microphone array beamformingabstractIn this paper we consider the problem of detecting speaker changes in audio signals recorded by distant microphones. It is shown that the possibility to exploit the spatial separation of speakers more than makes up the degradation in detection accuracy due to the increased source-to-sensor distance compared to close-talking microphones. Speaker direction information is derived from the filter coefficients of an adaptive Filter-and-Sum Beamformer and is combined with BIC analysis. The experimental results reveal significant improvements compared to BIC-only change detection, be it with the distant or close-talking microphone. Joerg Schmalenstroeer, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2006 | Uncertainty decoding for distributed speech recognition over error-prone networks
Valentin Ion, Reinhold Häb-Umbach |
Speech Commun. | 2 |
| 2005 | A Comparison of Soft-Feature Distributed Speech Recognition with Candidate Codecs for Speech Enabled Mobile ServicesabstractIn this paper we present a comparison of the recently proposed soft-feature distributed speech recognition (SFDSR) with the two evaluated candidate codecs for speech enabled services over wireless networks: adaptive multirate codec (AMR) and the ETSI extended advanced front-end for distributed speech recognition (XAFE). It is shown that SFDSR achieves the best recognition performance on a simulated GSM transmission, followed by XAFE and AMR. We also present some new results concerning SFDSR which demonstrate the versatility of the approach. Further, a simple method is introduced which considerably reduces the computational effort. Valentin Ion, Reinhold Häb-Umbach |
ICASSP (1) | 2 |
| 2005 | Acoustic filter-and-sum beamforming by adaptive principal component analysisabstractFor human-machine interfaces in distant-talking environments multichannel signal processing is often employed to obtain an enhanced signal for subsequent processing. In this paper we propose a novel adaptation algorithm for a filter-and-sum beamformer to adjust the coefficients of FIR filters to changing acoustic room impulses, e.g. due to speaker movement. A deterministic and a stochastic gradient ascent algorithm are derived from a constrained optimization problem, which iteratively estimates the eigenvector corresponding to the largest eigenvalue of the cross power spectral density of the microphone signals. The method does not require an explicit estimation of the speaker location. The experimental results show fast adaptation and excellent robustness of the proposed algorithm. Ernst Warsitz, Reinhold Häb-Umbach |
ICASSP (4) | 2 |
| 2005 | Speech processing in the networked home environment - a view on the amigo projectabstractFull interoperability of networked devices in the home has been kind of an elusive concept for quite some years. Amigo, an Integrated Project within the EU 6-th framework program, tries to make home networking a reality by addressing two key issues: First, it brings together many major players in the domestic appliances, communications, consumer electronics and computer industry to develop a common open source middleware platform. Second, emphasis is placed on the development of intelligent user services that make the benefit of a networked home environment tangible for the end user. This paper shows how speech processing can contribute to this second goal of user-friendly, personalized, context-aware services. 1. Reinhold Häb-Umbach, Basilis Kladis, Joerg Schmalenstroeer |
INTERSPEECH | 1 |
| 2005 | A comparison of particle filtering variants for speech feature enhancementabstractThis paper compares several particle filtering variants for speech feature enhancement in non-stationary noise environments. By analyzing the random processes of clean speech, noise and noisy speech, appropriate proposal densities are derived. The performances of the resulting particle filters, i.e. modified Sampling-Importance-Resampling (mod-SIR), auxiliary SIR and likelihood particle filter, are compared in terms of word accuracy achieved by the subsequent speech recognizer on the AURORA 2 database. It turns out that for the noises found in this database, noise compensation techniques that assume stationary noise work equally well. 1. Reinhold Häb-Umbach, Joerg Schmalenstroeer |
INTERSPEECH | 1 |
| 2005 | Unified probabilistic approach to error concealment for distributed speech recognitionabstractThe transmission errors in a wireless or packet oriented network may dramatically decrease the performance of a distributed speech recognition (DSR) system. Error concealment has been shown to be an effective way to mantain an acceptable word error rate when dealing with error prone communication channels. In this paper we propose an extension of our previously introduced soft features approach for the case that the soft-output of the channel decoder is not available at the server side of the DSR system. We found a simple method to estimate bit reliability information which still gives good speech recognition results. It is shown that some other error concealment schemes turn out to be special cases of the method proposed here. 1. Valentin Ion, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 2004 | Soft features for improved distributed speech recognition over wireless networksabstractA major drawback of distributed versus terminal-based speech recognition is the fact that transmission errors can lead to degraded recognition performance. In this paper we employ soft features to mitigate the effect of bit errors on wireless transmission links: At the receiver a posteriori probabilities of the transmitted feature vectors are computed by combining bit reliability information provided by the channel decoder and a priori knowledge about residual redundancy in the feature vectors. While the first-order moment of the a posteriori probability function is the MMSE estimate, the second-order moment is a measure of the uncertainty in the reconstructed features. We conducted realistic simulations of GSM transmission and achieved significant improvements in word accuracy compared to the error mitigation strategy described in the ETSI standard. Reinhold Häb-Umbach, Valentin Ion |
INTERSPEECH | 1 |
| 2004 | Adaptive beamforming combined with particle filtering for acoustic source localizationabstractWhile the main objective of adaptive Filter-and-Sum beamforming is to obtain an enhanced speech signal for subsequent processing like speech recognition, we show how speaker localization information can be derived from the filter coefficients. To increase localization accuracy, speaker tracking is performed by non-linear Bayesian state estimation, which is realized by sequential Monte Carlo methods. Improved acquisition and tracking performance was achieved even in highly reverberant environments, in comparison with both a Kalman Filter and a recently proposed Particle Filter operating on the output of a nonadaptive Delay-and-Sum beamformer. Reinhold Häb-Umbach, Sven Peschke, Ernst Warsitz |
INTERSPEECH | 1 |
| 2004 | Robust speaker direction estimation with particle filteringabstractThe paper is concerned with binaural signal processing for a bimodal human-robot interface with hearing and vision. The two microphone signals are processed to obtain an enhanced single-channel input signal for the subsequent speech recognizer and to localize the acoustic source, an important information for establishing a natural human-robot communication. We utilize a robust adaptive algorithm for filter-and-sum beamforming (FSB) and extract speaker direction information from the resulting FIR filter coefficients. Further, particle filtering is applied which conducts a nonlinear Bayesian tracking of speaker movement. Good location accuracy can be achieved even in highly reverberant environments. The results obtained outperform the conventional generalized cross correlation (GCC) method. Ernst Warsitz, Reinhold Häb-Umbach |
MMSP | 2 |
| 2002 | Employment of a multipath receiver structure in a combined GALILEO/UMTS receiverabstractCurrent navigation systems like GPS (Global Positioning System) and its Russian counterpart GLONASS (Global Navigation Satellite System) only evaluate the direct signal path. The receivers treat the reflected paths also reaching the receiver antenna as disturbance which has to be suppressed. Multipath affects the tracking accuracy by resulting in a degeneration of the S-curve of the DLL (delay locked loop). Nowadays the future European systems GALILEO and GPSIIF/III with two new signals are on the way to the market and it is time to think about new receiver structures. Therefore we investigated if it is possible to use multipath for navigation constructively. Renke Bischoff, Reinhold Häb-Umbach, Wolfgang Schulz 0003, Guenter Heinrichs |
VTC Spring | 2 |
| 2002 | Large vocabulary continuous speech recognition of Broadcast News - The Philips/RWTH approach
Peter Beyerlein, Xavier L. Aubert, Reinhold Häb-Umbach, Matthew Harris, Dietrich Klakow, Andreas Wendemuth, Sirko Molau, Hermann Ney, Michael Pitz, Achim Sixtus |
Speech Commun. | 3 |
| 2001 | Multiclass Linear Dimension Reduction by Weighted Pairwise Fisher CriteriaabstractWe derive a class of computationally inexpensive linear dimension reduction criteria by introducing a weighted variant of the well-known K-class Fisher criterion associated with linear discriminant analysis (LDA). It can be seen that LDA weights contributions of individual class pairs according to the Euclidean distance of the respective class means. We generalize upon LDA by introducing a different weighting function. Marco Loog, Robert P. W. Duin, Reinhold Häb-Umbach |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2001 | Automatic generation of phonetic regression class trees for MLLR adaptationabstractIn this paper, it is shown that a correlation criterion is the appropriate criterion for bottom-up clustering to obtain broad phonetic class regression trees for maximum likelihood linear regression (MLLR)-based speaker adaptation. The correlation structure among speech units is estimated on the speaker-independent training data. In adaptation experiments the tree outperformed a regression tree obtained from clustering according to closeness in acoustic space and achieved results comparable with those of a manually designed broad phonetic class tree. Reinhold Häb-Umbach |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | LDA derived cepstral trajectory filters in adverse environmental conditionsabstractAmongst several data driven approaches for designing filters for the time sequence of spectral parameters, the linear discriminant analysis (LDA) based method has been proposed for automatic speech recognition. Here we apply LDA-based filter design to cepstral features, which better match the inherent assumption of this method that feature vector components are uncorrelated. Extensive recognition experiments have been conducted both on the standard TIMIT phone recognition task and on a proprietary 130-words command word task under various adverse environmental conditions, including reverberant data with real-life room impulse responses and data processed by acoustic echo cancellation algorithms. Significant error rate reductions have been achieved when applying the novel long-range feature filters compared to standard approaches employing cepstral mean normalization and delta and delta-delta features, in particular when facing acoustic echo cancellation scenarios and room reverberation. For example, the phone accuracy on reverberated TIMIT data could be increased from 50.7% to 56.0%. Markus Lieb, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2000 | Multi-Class Linear Feature Extraction by Nonlinear PCAabstractThe traditional way to find a linear solution to feature extraction problems is based on the maximization of the class-between scatter over the class-within scatter (Fisher's mapping). For the multi-class problem this is sub-optimal due to class conjunctions, even for the simple situation of normal distributed classes with identical covariance matrices. We propose a novel, equally fast method, based on nonlinear principal component analysis (PCA). Although still sub-optimal, it may avoid the class conjunction. The proposed method is experimentally compared with Fisher's mapping and with a neural network based approach to nonlinear PCA. It appears to outperform the both methods. Robert P. W. Duin, Marco Loog, Reinhold Häb-Umbach |
ICPR | 3 |
| 2000 | Data-driven phonetic regression class tree estimation for MLLR adaptation
Reinhold Häb-Umbach |
INTERSPEECH | 1 |
| 2000 | Multi-class linear dimension reduction by generalized Fisher criteriaabstractLinear Disciminant Analysis is in general unable to find the lower-dimensional feature space which maximizes the class discrimination, even if the class distributions can be assumed to be very simple, e.g. Gaussians with identical covariance matrices. In this paper we reformulate the-class Fisher criterion as a sum of �-class Fisher criteria. This formulation allows to weigh class pair contributions according to their relevance for classification. Further it offers an obvious way how to cope with heteroscedastic models. We propose a particular weighting scheme which attempts to approximate the pairwise Bayes error. Moderate improvements are obtained on the TIMIT phoneme classification task. 1. Marco Loog, Reinhold Häb-Umbach |
INTERSPEECH | 2 |
| 1999 | Investigations on inter-speaker variability in the feature spaceabstractWe apply Fisher variate analysis to measure the effectiveness of speaker normalization techniques. A trace criterion, which measures the ratio of the variations due to different phonemes compared to variations due to different speakers, serves as a first assessment of a feature set without the need for recognition experiments. By using this measure and by recognition experiments we demonstrate that cepstral mean normalization also has a speaker normalization effect, in addition to the well-known channel normalization effect. Similarly vocal tract normalization (VTN) is shown to remove inter-speaker variability. For VTN we show that normalization on a per sentence basis performs better than normalization on a per speaker basis. Recognition results are given on Wall Street Journal and Hub-4 databases. Reinhold Häb-Umbach |
ICASSP | 1 |
| 1999 | The philips/RWTH system for transcription of broadcast newsabstractThis paper contains a description of the Philips/RWTH 1998 HUB4 system which has been build in a joint eort of Philips Research Laboratories Aachen and Aachen University o f T echnology.We will focus our discussion on recent improvements compared to the original 1997 HUB4 system and evaluate them on the HUB4'97 evaluation data.The paper will deal with 1. a rough system overview including feature extraction, acoustic training, audio stream segmentation, and decoding 2. log-linear interpolation of distance-language models, 3. and the integration of various acoustic and language models via Discriminative Model Combination (DMC).The performance of the described system is 23% (relative) better than the performance of the 1997 Philips HUB4 system.A w ord error rate of 17.9% was achieved on the 1997 HUB4 evaluation set, compared to 23.5% using the original 1997 system. Peter Beyerlein, Xavier L. Aubert, Reinhold Häb-Umbach, Matthew Harris, Dietrich Klakow, Andreas Wendemuth, Sirko Molau, Michael Pitz, Achim Sixtus |
EUROSPEECH | 3 |
| 1999 | An investigation of cepstral parameterisations for large vocabulary speech recognitionabstractWe examined variants of MFCC and PLP cepstral parameterisations in the context of large vocabulary continuous speech recognition under different acous-tical environmental conditions: Compared to MFCC, mel-frequency PLP uses a cubic root intensity-to-loudness law, and an LPC analysis is applied to the mel-warped spectrum. In LPC-smoothed MFCC, the only difference to MFCC is the additional LPC smoothing of the warped spectrum. While neither technique was able to significantly outperform the MFCC parameterisation in our setup which includes an LDA feature transformation, feature set combination via DMC at the acoustic likelihood level and via ROVER at the recognized word level delivered small but consistent improvements. Reinhold Häb-Umbach, Marco Loog |
EUROSPEECH | 1 |
| 1999 | A study of broadcast news audio stream segmentation and segment clusteringabstractThe paper deals with some aspects of production and perception of yes-no questions and non-final clauses in Dutch. The analysis of the prosodic characteristics of naturally produced utterances and the subjects’ responses obtained in the course a series of perception experiments with naturally produced and modified utterances yielded the following conclusions: • the identification of utterance type is to the largest extent determined by the presence of relevant prosodic cues in the nuclear accent • cues to utterance type in the pre-nuclear pattern appear to be used by listeners when the nuclear accent type allows of several possible interpretations • timing of the nuclear accent may be considered a relevant cue to utterance type, «early» vs. «late» rise accounting for «non-final» vs. «interrogative» preferences, «early» vs. «late» fall eliciting «statement» vs. «question/statement» reactions. INTRODUCTION The establishing of a set of categorically distinct intonation units forming the intonation system of a language is still one of the most challenging tasks for linguists. It is not yet clear how exactly the linguistic relevance of a certain type of intonation unit should be tested and some disagreement exists as to what differences should be called categorical rather than gradient [6]. There is no agreement either as to the function of intonation. Are intonation units associated with certain invariant categories of meaning and if so, in what terms should these meaningful distinctions be described? These and other fundamental questions are still open for discussion. At the same time it is evident that intonation is capable of changing the interpretation of an utterance and thus may be said to perform a certain communicative function, which finds its expression, among others, in the specific patterning of various utterance types. Previous research in the intonation of Dutch questions (e.g. [2], [3], [4]) has proved the already reported for other languages founding that the concept of interrogativity is mainly associated with a local terminal rise. Other cues to interrogativity include a raised register and absence of downtrend. Non-finality (or «continuation» we shall leave the terminological discussion beyond the scope of the present work) is also said to be cued by rising pitch. In her work devoted to the melodic marking of continuation versus questions in Dutch [2] J. Caspers claims that while the former are signalled by the accent-lending rise followed by sustained level pitch (1O in the IPO notation [1]), the latter are associated with a combination of an accent-lending rise and a final rise (12). One possible interpretation of these findings is that they provide extra evidence in support of the previously mentioned claim that questions are characterised by higher pitch values than other utterance types. The study reported here is concerned with finding further prosodic cues to non-finality and interrogativity. We limited our research exclusively to pitch characteristics (being fully aware of the communicative importance of other classes of prosodic features). The questions addressed were the following: • What are the most typical intonation patterns realised in non-final clauses and yes-no questions (in terms of nuclear and pre-nuclear patterns and their combinations) • To what extent can the perception of utterance type be influenced by the cues contained in the prenuclear vs. nuclear part of the utterance? • What is the role played by the timing of the nuclear accent in the perception of the utterance type ( Rise 1 vs. 2 and Fall A vs. (1)A in the IPO notation)? EXPERIMENTS 1: ANALYSIS AND PERCEPTION OF (FRAGMENTS OF) NATURALLY PRODUCED UTTERANCES 9 Yes-no questions with the inverted word order and 9 lexically and syntactically identical to them non-final clauses (mostly subordinate clauses of condition and concession also having the inverted word order) were embedded in appropriate contexts, printed on cards and read by 8 native speakers of Dutch (5 male, 3 female). The acquired recordings (144 utterances) were digitised and F0 contours were extracted using speech analysis programs WinCECIL (distributed by SIL as freeware) and PRAAT (developed by P. Boersma and D. Weenink). Pitch patterns produced by the speakers were transcribed using largely the notation developed in the framework of the Dutch School [1], with some modifications introduced in the course of the experimental analysis of the data, and then classified (following the British tradition in intonation analysis [7]) according to the type of nuclear and pre-nuclear pattern and frequency of occurrence in either type of utterance. 6 types of nuclear pitch accents were distinguished («&» means that the pitch movements are associated with one accented syllable, A* is used to indicate a delayed fall not distinguished in the IPO system): IPO Notation: Rise 1 (early ) 1 Rise 2 (late) (a)2, (b)A&2 Matthew Harris, Xavier L. Aubert, Reinhold Häb-Umbach, Peter Beyerlein |
EUROSPEECH | 3 |
| 1998 | A study on speaker normalization using vocal tract normalization and speaker adaptive trainingabstractAlthough speaker normalization is attempted in very different manners, vocal tract normalization (VTN) and speaker adaptive training (SAT) share many common properties. We show that both lead to more compact representations of the phonetically relevant variations of the training data and that both achieve improved error rate performance only if a complementary normalization or adaptation operation is conducted on the test data. Algorithms for fast test speaker enrolment are presented for both normalization methods: in the framework of SAT, a pre-transformation step is proposed, which alone, i.e. without subsequent unsupervised MLLR adaptation, reduces the error rate by almost 10% on the WSJ 5k test sets. For VTN, the use of a Gaussian mixture model makes obsolete a first recognition pass to obtain a preliminary transcription of the test utterance at hardly any loss in performance. Lutz Welling, Reinhold Häb-Umbach, X. Zubert, N. Haberland |
ICASSP | 2 |
| 1997 | Signal representations for hidden Markov model based online handwriting recognitionabstractAddresses the problem of online, writer-independent, unconstrained handwriting recognition. Based on hidden Markov models (HMM), which are successfully employed in speech recognition tasks, we focus on representations which address scalability, recognition performance and compactness. 'Delayed' features are introduced which integrate more global, handwriting specific knowledge into the HMM representation. These features lead to larger error-rate reduction than 'delta' features which are known from speech recognition and even require fewer additional components. Scalability is addressed with a size-independent representation. Compactness is achieved with linear discriminant analysis. The representations are discussed and the results for a mixed-style word recognition task with vocabularies of 200 (up to 99% correct words) and 20000 words (up to 88.8% correct words) are given. Hans J. G. A. Dolfing, Reinhold Häb-Umbach |
ICASSP | 2 |
| 1997 | European speech databases for telephone applicationsabstractThe SpeechDat project aims to produce speech databases for all official languages of the European Union and some major dialectal variants and minority languages resulting in 28 speech databases. They will be recorded over fixed and mobile telephone networks. This will provide a realistic basis for training and assessment of both isolated and continuous-speech utterances, employing whole-word or subword approaches, and thus can be used for developing voice driven teleservices including speaker verification. The specification of the databases has been developed jointly, and is essentially the same for each language to facilitate dissemination and use. There will be a controlled variation among the speakers concerning sex, age, dialect, environment of call, etc. The validation of all databases will be carried out centrally. The SpeechDat databases will be transferred to ELRA for distribution. The next databases to be recorded will cover East European languages. Harald Höge, Herbert S. Tropf, Richard Winski, Henk van den Heuvel, Reinhold Häb-Umbach, Khalid Choukri |
ICASSP | 5 |
| 1997 | Robust speech recognition for wireless networks and mobile telephonyabstractThe increased popularity of mobile telephony introduces both challenges and opportunitites for automatic speech recognition. ASR offers ways to simplify the use of mobile phones, notably in hands- and eyes-busy situations. However, the acoustic environment can be severely degraded and the wireless network may add additional distortions to the speech signal. This paper gives an overview of the sources of degradation and attempts to robust speech recognition for mobile communications. Emphasis is placed on approaches which are suitable for implementation in mobile terminals. Two example applications are described which illustrate the robustness issues and design considerations typical of low-cost noisy speech recognition: voice-dialling in a GSM phone and hands-free digit recognition in the car. Reinhold Häb-Umbach |
EUROSPEECH | 1 |
| 1997 | Acoustic front ends for speaker-independent digit recognition in car environmentsabstractThis paper describes speaker-independent speech recognition experiments concerning acoustic front end processing on a speech database that was recorded in 3 different cars. We investigate different feature analysis approaches (mel-filter bank, mel-cepstrum, perceptually linear predictive coding) and present results with noise compensation techniques based on spectral subtraction. Although the methods employed lead to considerable error rate reduction the error analysis shows that low signal-to-noise ratios are still a problem Detlev Langmann, Alexander Fischer, Friedhelm Wuppermann, Reinhold Häb-Umbach, Thomas Eisele |
EUROSPEECH | 4 |
| 1997 | The development of a command-based speech interface for a telephone answering machine
Stephan Gamm, Reinhold Häb-Umbach, Detlev Langmann |
Speech Commun. | 2 |
| 1996 | A comparative study of linear feature transformation techniques for automatic speech recognitionabstractAlthough widely used, there are still open questions concerning which properties of Linear Discriminant Analysis (LDA) do account for its success in many speech recognition systems.In order to gain more insight into the nature of the transformation we compare LDA with mel-cepstral feature vectors with respect to the following criteria: decorrelation and ordering property, invariance under linear transforms, automatic learning of dynamical features, and data dependence of the transformation. Thomas Eisele, Reinhold Häb-Umbach, Detlev Langmann |
ICSLP | 2 |
| 1996 | FRESCO: the French telephone speech data collection - part of the european Speechdat(m) project
Detlev Langmann, Reinhold Häb-Umbach, Lou Boves, Els den Os |
ICSLP | 2 |
| 1995 | Application of clustering techniques to mixture density modelling for continuous-speech recognitionabstractClustering techniques have been integrated at different levels into the training procedure of a continuous-density hidden Markov model (HMM) speech recognizer. These clustering techniques can be used in two ways. First acoustically similar states are tied together. It will help to reduce the number of parameters but also allow to train otherwise rarely seen states together with more robust ones (state-tying). Secondly densities are clustered across states, this reduces the number of densities while at the same time keeping the best performances of our recognizer (density-clustering). We have applied these techniques both to word-based small-vocabulary and phoneme-based large-vocabulary recognition tasks. On the WSJ task, we could achieve a reduction of the word error rate by 7%. On the TI/NIST-connected digit task, the number of parameters was reduced by a factor 2-3 while keeping the same string error rate. Christian Dugast, Peter Beyerlein, Reinhold Häb-Umbach |
ICASSP | 3 |
| 1995 | Automatic transcription of unknown words in a speech recognition systemabstractWe address the problem of automatically finding an acoustic representation (i.e. a transcription) of unknown words as a sequence of subword units, given a few sample utterances of the unknown words, and an inventory of speaker-independent subword units. The problem arises if a user wants to add his own vocabulary to a speaker-independent recognition system simply by speaking the words a few times. Two methods are investigated which are both based on a maximum-likelihood formulation of the problem. The experimental results show that both automatic transcription methods provide a good estimate of the acoustic models of unknown words. The recognition error rates obtained with such models in a speaker-independent recognition task are clearly better than those resulting from separate whole-word models. They are comparable with the performance of transcriptions drawn from a dictionary. Reinhold Häb-Umbach, Peter Beyerlein, Eric Thelen |
ICASSP | 1 |
| 1995 | Human factors of a voice-controlled car stereo
Reinhold Häb-Umbach, Stephan Gamm |
EUROSPEECH | 1 |
| 1995 | Continuous speech dictation - From theory to practice
Volker Steinbiss, Hermann Ney, Ute Essen, Bach-Hiep Tran, Xavier L. Aubert, Christian Dugast, Reinhard Kneser, Hans-Günter Meier, Martin Oerder, Reinhold Häb-Umbach, Dieter Geller, W. Höllerbauer, H. Bartosik |
Speech Commun. | 10 |
| 1994 | An Overview of the Philips Research System for Large Vocabulary Continuous Speech RecognitionabstractThis paper gives an overview of a research system for phoneme based, large vocabulary continuous speech recognition. The system to be described has been applied to the SPICOS task, the DARPA RM task and a 12000 word dictation task. Experimental results for these three tasks will be presented. Like many other systems, the recognition architecture is based on an integrated statistical approach. In this paper, we describe the characteristic features of the system as opposed to other systems: (1) The Viterbi criterion is consistently applied both in training and testing. (2) Continuous mixture densities are used without any tying or smoothing; this approach can be viewed as a sort of ‘statistical template matching’. (3) Time-synchronous beam search is used consistently throughout all tasks; extensions using a tree organization of the vocabulary and phoneme lookahead are presented so that a 12000 word task can be handled. Hermann Ney, Volker Steinbiss, Reinhold Häb-Umbach, Bach-Hiep Tran, Ute Essen |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1994 | Improvements in beam search for 10000-word continuous-speech recognitionabstractThe authors describe the improvements in a time-synchronous beam search strategy for a 10000-word continuous-speech recognition task. Basically they introduced two measures, namely a tree organization of the pronunciation lexicon and a novel look-ahead technique at the phoneme level. The experimental tests performed showed that the number of state hypotheses could be reduced from 50000 to 3000, i.e., by a factor of about 17. At the same time, the word error rate did not increase.> Reinhold Häb-Umbach, Hermann Ney |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | Continuous mixture densities and linear discriminant analysis for improved context-dependent acoustic models
Xavier L. Aubert, Reinhold Häb-Umbach, Hermann Ney |
ICASSP (2) | 2 |
| 1993 | Improvements in connected digit recognition using linear discriminant analysis and mixture densities
Reinhold Häb-Umbach, Dieter Geller, Hermann Ney |
ICASSP (2) | 1 |
| 1993 | The Philips research system for large-vocabulary continuous-speech recognitionabstractThis paper gives a status report of the Philips research system for phoneme-based, large-vocabulary, continuous-speech recognition. Like for many other systems, the recognition architecture is based on an integrated statistical approach. We describe the characteristic features of the system as opposed to other systems: 1. The Viterbi criterion is consistently applied both in training and testing. 2. Continuous mixture densities are used without tying or smoothing. 3. Time-synchronous beam search in connection with a phoneme look-ahead is applied to a tree-organized lexicon. The system has been successfully applied to the American English DARPA RM task. Here, we report experimental results for a German 13 000-word Philips internal dictation task. In addition to the scientific prototype, a PC version has been set up which is described here for the first time. Volker Steinbiss, Hermann Ney, Reinhold Häb-Umbach, B.-H. Iran, Ute Essen, Reinhard Kneser, Martin Oerder, Hans-Günter Meier, Xavier L. Aubert, Christian Dugast, Dieter Geller, W. Höllerbauer, H. Bartosik |
EUROSPEECH | 3 |
| 1993 | Design and use of speech recognition algorithms for a mobile radio telephone
Stefan Dobler, Dieter Geller, Reinhold Häb-Umbach, Peter Meyer, Hermann Ney, Hans-Wilhelm Rühl |
Speech Commun. | 3 |
| 1992 | Linear discriminant analysis for improved large vocabulary continuous speech recognitionabstractThe interaction of linear discriminant analysis (LDA) and a modeling approach using continuous Laplacian mixture density HMM is studied experimentally. The largest improvements in speech recognition could be obtained when the classes for the LDA transform were defined to be sub-phone units. On a 12000 word German recognition task with small overlap between training and test vocabulary a reduction in error rate by one-fifth was achieved compared to the case without LDA. On the development set of the DARPA RM1 task the error rate was reduced by one-third. For the DARPA speaker-dependent no-grammar case, the error rate averaged over 12 speakers was 9.9%. This was achieved with a recognizer using LDA and a set of only 47 Viterbi-trained context-independent phonemes.> Reinhold Häb-Umbach, Hermann Ney |
ICASSP | 1 |
| 1992 | Improvements in beam search for 10000-word continuous speech recognitionabstractThe author describes the improvements in a time synchronous beam search strategy for a 10000-word continuous speech recognition task. The improvements are based on two measures: a tree-organization of the pronunciation lexicon and a novel look-ahead technique at the phoneme level, both of which interact directly with the detailed search at the state levels of the phoneme models. Experimental tests were performed for four speakers on a 12306-word task. As a result of the above measures, the overall search effort was reduced by a factor of 17 without a loss in recognition accuracy.> Hermann Ney, Reinhold Häb-Umbach, Bach-Hiep Tran, Martin Oerder |
ICASSP | 2 |
| 1992 | Trellis Codes for Partial-Response Magnetooptical Direct Overwrite RecordingabstractThe authors present conditions on the error sequences between channel input sequences which guarantee certain lower bounds on the free Euclidian distance at the output of a partial-response (PR) class I or II channel. From these expressions, trellis codes are derived which improve performance of binary signaling over noisy PR channels with reduced complexity maximum-likelihood sequence detection. They are shown to be compatible with the input restriction caused by the magnetooptical resonant coil direct overwrite recording scheme. The codes achieve high signal-to-noise ratio coding gains of 3 dB (on PR class I) and 2.2 dB (on PR class II) with rates as close to, but strictly less than, the capacity of the initial input restriction as desired. The performance of these codes is analyzed with an optical channel simulation system which shows that one code has the rare but highly desirable property that its maximum-likelihood sequence detector (MLSD) is less complex than the MLSD of the reference system and still achieves an error rate performance gain of 1.8 dB.> Reinhold Häb-Umbach, Robert T. Lynch Jr. |
IEEE J. Sel. Areas Commun. | 1 |
| 1992 | A modified trellis coding technique for partial response channelsabstractThe problem of trellis coding for multilevel baseband transmission over partial response channels with transfer polynomials of the form (1+or-D/sup N/) is addressed. The novel method presented here accounts for the channel memory by using multidimensional signal sets and partitioning the signal set present at the noiseless channel output. It is shown that this coding technique can be viewed as a generalization of a well-known procedure for binary signaling: the concatenation of convolutional codes and inner block codes that are tuned to the channel polynomial. It results in high coding gains with moderate complexity if some bandwidth expansion is accepted.> Reinhold Häb-Umbach |
IEEE Trans. Commun. | 1 |
| 1991 | A look-ahead search technique for large vocabulary continuous speech recognitionabstractIn a large vocabulary continuous speech recognition task the search for the best (in the maximum-a-posteriori sense) word sequence is the most (computing) time consuming part of the system. End-of-word hypotheses are created almost every time frame. With a stochastic language model every lexicon entry is an admissible successor candidate. By using a module which scores the word candidates according to their acoustic feasibility ahead of the current time frame, the search cost can be considerably reduced. Only the fraction of the words with favourable fast match scores will be further processed in the detailed match, where the likelihood of a segment of acoustics given the word model is computed. We derive a novel word selection strategy which is consistent in the sense that it introduces no additional decoding errors and which still reduces the search space by a factor of 2 - 3 compared to standard Viterbi beam search. Giving up the consistency requirement, pruning strategies can be deduced which further reduce the search effort significantly: the size of the word startup list is reduced to 2% - 4% of its original size with a modest increase in error rate by l%-2%. Reinhold Häb-Umbach, Hermann Ney |
EUROSPEECH | 1 |
| 1989 | A systematic approach to carrier recovery and detection of digitally phase modulated signals of fading channelsabstractThe problem of optimal carrier recovery and detection of digitally phase modulated signals on fading channels by using a nonstructured approach is presented, i.e. no constraint is placed on the receiver structure. First, the optimal receiver is derived for digitally phase-modulated signals when transmitted over a frequency-nonselective fading channel with memory. The memory results from the fact that usually the coherence time of the channel is larger than the symbol period. Symbols adjacent in time cannot be detected independently and therefore the well-known quadratic receiver is not optimal in this case. A maximum a posteriori (MAP) detector is derived and explicitly utilizes the channel memory for carrier recovery. The derivation shows that the optimal carrier recovery is, under certain conditions, a Kalman filter. Some attractive properties of this carrier recovery unit (including the absence of hang up) are discussed. Then the error rate of several digital modulation schemes is calculated taking the performance of the filter into account. The differences in susceptibility of the modulation schemes to carrier phase jitter are specified.> Reinhold Häb-Umbach, Heinrich Meyr |
IEEE Trans. Commun. | 1 |