VLDB 2026 Research / reviewers in the wild / expert
Timo Gerkmann
dblp:57/44
· DBLP profile ↗
100ranked-venue papers
8as first author
53since 2021 · last 2026
0000-0002-8678-4699ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 72 · 6 first-author · 39 since 2021Artificial intelligence and machine learning · 50 · 3 first-author · 29 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamics of Collective Group Affect: Group-Level Annotations and the Multimodal Modeling of Convergence and DivergenceabstractCollaborating in a purposive group, whether face-to-face or virtually, involves continuously expressing emotions and interpreting those of other group members. As such, understanding group affect is essential to comprehending how groups interact and succeed in collaborative efforts. In this study, we move beyond individual-level affect and investigate group-level affect—a collective phenomenon that reflects the shared mood or emotions among group members at a particular moment. As the first in the literature, we gather annotations for group-level affective expressions in purposive group interactions using a fine-grained temporal approach (15 second windows) that also captures the inherent dynamics of this collective construct. To this end, we extensively train annotators and develop an annotation procedure specifically tuned to capture the entire scope of the group interaction from one interaction moment to the next. In addition, we model the ebb and flow of group affect by accounting for the underlying convergence (driven by emotional contagion) and divergence (resulting from emotional reactivity) of affective expressions among group members. To capture these interpersonal dynamics, we employ two approaches: (i) extracting synchrony-based handcrafted features from both audio and visual modalities, and (ii) introducing a novel, data-driven graph neural network to model interpersonal dynamics among group members. Our results highlight the advantages of the graph network over the handcrafted features in modeling group affect, while also emphasizing the importance of temporal modeling and incorporating multimodal cues. Additionally, our analysis of affective convergence and divergence reveals that groups tend to diverge in their social signals during neutral collective affect, while exhibiting convergence during more emotionally intense moments. These insights are drawn from comparative results across both modeling techniques. Navin Raj Prabhu, Maria Tsfasman, Catharine Oertel, Timo Gerkmann, Nale Lehmann-Willenbrock |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration ModelingabstractSpeech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic properties such as fundamental frequency, spectral envelope and energy, but often lack the ability to control the duration of sounds. To address this, we propose a duration modeling framework using resynthesis-based discrete content representations, enabling modification of speech duration to reflect target emotions and achieve controllable speech rates without using parallel data. Experimental results reveal that the inclusion of the proposed duration modeling framework significantly enhances emotional expressiveness, in the in-the-wild MSP-Podcast dataset. Analyses show that low-arousal emotions correlate with longer durations and slower speech rates, while high-arousal emotions produce shorter, faster speech. Navin Raj Prabhu, Danilo de Oliveira, Nale Lehmann-Willenbrock, Timo Gerkmann |
ASRU | 4 |
| 2025 | Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask EstimationabstractDue to their robustness and flexibility, neural-driven beamformers are a popular choice for speech separation in challenging environments with a varying amount of simultaneous speakers alongside noise and reverberation. Time-frequency masks and relative directions of the speakers regarding a fixed spatial grid can be used to estimate the beamformer’s parameters. To some degree, speaker-independence is achieved by ensuring a greater amount of spatial partitions than speech sources. In this work, we analyze how to encode both mask and positioning into such a grid to enable joint estimation of both quantities. We propose mask-weighted spatial likelihood coding and show that it achieves considerable performance in both tasks compared to baseline encodings optimized for either localization or mask estimation. In the same setup, we demonstrate superiority for joint estimation of both quantities. Conclusively, we propose a universal approach which can replace an upstream sound source localization system solely by adapting the training framework, making it highly relevant in performance-critical scenarios. Jakob Kienegger, Alina Mannanova, Timo Gerkmann |
ICASSP | 3 |
| 2025 | Investigating Training Objectives for Generative Speech EnhancementabstractGenerative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schrodinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schrodinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain1. Julius Richter, Danilo de Oliveira, Timo Gerkmann |
ICASSP | 3 |
| 2025 | HRTF Estimation using a Score-based PriorabstractWe present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human speech. The impulse response of the room is estimated along with the HRTF by optimizing a parametric model of reverberation based on the statistical behaviour of room acoustics. The posterior distribution of HRTF given the reverberant measurement and excitation signal is modelled using the score-based HRTF prior and a log-likelihood approximation. We show that the resulting method outperforms several baselines, including an oracle recommender system that assigns the optimal HRTF in our training set based on the smallest distance to the true HRTF at the given direction of arrival. In particular, we show that the diffusion prior can account for the large variability of high-frequency content in HRTFs. Etienne Thuillier, Jean-Marie Lemercier, Eloi Moliner, Timo Gerkmann, Vesa Välimäki |
ICASSP | 4 |
| 2025 | FlowDec: A flow-based full-band general audio codec with high perceptual qualityabstractWe propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from speech to general audio and move from 24 kbit/s to as low as 4 kbit/s, while improving output quality and reducing the required postfilter DNN evaluations from 60 to 6 without any fine-tuning or distillation techniques. We provide theoretical insights and geometric intuitions for our approach in comparison to ScoreDec as well as another recent work that uses flow matching, and conduct ablation studies on our proposed components. We show that FlowDec is a competitive alternative to the recent GAN-dominated stream of neural codecs, achieving FAD scores better than those of the established GAN-based codec DAC and listening test scores that are on par, and producing qualitatively more natural reconstructions for speech and harmonic structures in music. Simon Welker, Matt Le 0001, Ricky T. Q. Chen, Wei-Ning Hsu, Timo Gerkmann, Alexander Richard, Yi-Chiao Wu |
ICLR | 5 |
| 2025 | Steering Deep Non-Linear Spatially Selective Filters for Weakly Guided Extraction of Moving Speakers in Dynamic Scenarios
Jakob Kienegger, Timo Gerkmann |
INTERSPEECH | 2 |
| 2025 | Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency
Bunlong Lay, Rostilav Makarov, Timo Gerkmann |
INTERSPEECH | 3 |
| 2025 | Real-Time Diffusion Buffer for Speech Enhancement On A Laptop
Bunlong Lay, Rostilav Makarov, Timo Gerkmann |
INTERSPEECH | 3 |
| 2025 | Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
Danilo de Oliveira, Julius Richter, Jean-Marie Lemercier, Simon Welker, Timo Gerkmann |
INTERSPEECH | 5 |
| 2024 | Single and Few-Step Diffusion for Generative Speech EnhancementabstractDiffusion models have shown promising results in single-channel speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score estimation is called multiple times to solve the iterative reverse process. This results in a slow inference process and causes discretization errors that accumulate over the sampling trajectory. In this paper, we address these limitations through a two-stage training approach. In the first stage, we train the diffusion model the usual way using the generative denoising score matching loss. In the second stage, we compute the enhanced signal by solving the reverse process and compare the resulting estimate to the clean speech target using a predictive loss. We show that using this second training stage enables achieving the same performance as the baseline model using only 5 function evaluations instead of 60 function evaluations. While the performance of usual generative diffusion algorithms drops dramatically when lowering the number of function evaluations to obtain single-step diffusion, we show that our proposed method keeps a steady performance and therefore largely outperforms the diffusion baseline in this setting and also generalizes better than its predictive counterpart1. Bunlong Lay, Jean-Marie Lemercier, Julius Richter, Timo Gerkmann |
ICASSP | 4 |
| 2024 | Distilling Hubert with LSTMs via Decoupled Knowledge DistillationabstractMuch research effort is being applied to the task of compressing the knowledge of self-supervised models, which are powerful, yet large and memory consuming. Existing works generally distill internal features of self-supervised Transformer models. In this work, aiming at more flexibility in the design of the student model, we apply the method of Knowledge Distillation and its more recently proposed extension, Decoupled Knowledge Distillation, to the task of distilling HuBERT. We achieve this by leveraging the cluster prediction pre-training task of HuBERT, which provides valuable targets for the distillation objective. We thus propose to exploit the acquired flexibility to distill HuBERT’s Transformer layers into an LSTM-based model that reduces the number of parameters even below DistilHuBERT and at the same time shows improved performance in automatic speech recognition. Danilo de Oliveira, Timo Gerkmann |
ICASSP | 2 |
| 2024 | A Flexible Online Framework for Projection-Based Stft Phase RetrievalabstractSeveral recent contributions in the field of iterative STFT phase retrieval have demonstrated that the performance of the classical Griffin-Lim method can be considerably improved upon. By using the same projection operators as Griffin-Lim, but combining them in innovative ways, these approaches achieve better results in terms of both reconstruction quality and required number of iterations, while retaining a similar computational complexity per iteration. However, like GriffinLim, these algorithms operate in an offline manner and thus require an entire spectrogram as input, which is an unrealistic requirement for many real-world speech communication applications. We propose to extend RTISI—an existing online (frame-by-frame) variant of the Griffin-Lim algorithm—into a flexible framework that enables straightforward online implementation of any algorithm based on iterative projections. We further employ this framework to implement online variants of the fast Griffin-Lim algorithm, the accelerated Griffin-Lim algorithm, and two algorithms from the optics domain. Evaluation results on speech signals show that, similarly to the offline case, these algorithms can achieve a considerable performance gain compared to RTISI. Tal Peer, Simon Welker, Johannes Kolhoff, Timo Gerkmann |
ICASSP | 4 |
| 2024 | EMOCONV-Diff: Diffusion-Based Speech Emotion Conversion for Non-Parallel and in-the-Wild DataabstractSpeech emotion conversion is the task of converting the expressed emotion of a spoken utterance to a target emotion while preserving the lexical content and speaker identity. While most existing works in speech emotion conversion rely on acted-out datasets and parallel data samples, in this work we specifically focus on more challenging in-the-wild scenarios and do not rely on parallel data. To this end, we propose a diffusion-based generative model for speech emotion conversion, the EmoConv-Diff, that is trained to reconstruct an input utterance while also conditioning on its emotion. Subsequently, at inference, a target emotion embedding is employed to convert the emotion of the input utterance to the given target emotion. As opposed to performing emotion conversion on categorical representations, we use a continuous arousal dimension to represent emotions while also achieving intensity control. We validate the proposed methodology on a large in-the-wild dataset, the MSP-Podcast v1.10. Our results show that the proposed diffusion model is indeed capable of synthesizing speech with a controllable target emotion. Crucially, the proposed approach shows improved performance along the extreme values of arousal and thereby addresses a common challenge in the speech emotion conversion literature. Navin Raj Prabhu, Bunlong Lay, Simon Welker, Nale Lehmann-Willenbrock, Timo Gerkmann |
ICASSP | 5 |
| 2024 | Live Iterative Ptychography with Projection-Based AlgorithmsabstractIn this work, we demonstrate that the ptychographic phase problem can be solved in a live fashion during scanning, while data is still being collected. We propose a generally applicable modification of the widespread projection-based algorithms such as Error Reduction (ER) and Difference Map (DM). This novel variant of ptychographic phase retrieval enables immediate visual feedback during experiments, reconstruction of arbitrary-sized objects with a fixed amount of computational resources, and adaptive scanning. By building upon the Real-Time Iterative Spectrogram Inversion (RTISI) family of algorithms from the audio processing literature, we show that live variants of projection-based methods such as DM can be derived naturally and may even achieve higher-quality reconstructions than their classic non-live counterparts with comparable effective computational load. Simon Welker, Tal Peer, Henry N. Chapman, Timo Gerkmann |
ICASSP | 4 |
| 2024 | An Analysis of the Variance of Diffusion-based Speech Enhancement
Bunlong Lay, Timo Gerkmann |
INTERSPEECH | 2 |
| 2024 | The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
Danilo de Oliveira, Simon Welker, Julius Richter, Timo Gerkmann |
INTERSPEECH | 4 |
| 2024 | EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker, Bunlong Lay, Shinji Watanabe 0001, Alexander Richard, Timo Gerkmann |
INTERSPEECH | 8 |
| 2024 | End-to-End Label Uncertainty Modeling in Speech Emotion Recognition Using Bayesian Neural Networks and Label Distribution LearningabstractTo train machine learning algorithms to predict emotional expressions in terms of arousal and valence, annotated datasets are needed. However, as different people perceive others' emotional expressions differently, their annotations are subjective. To account for this, annotations are typically collected from multiple annotators and averaged to obtain ground-truth labels. However, when exclusively trained on this averaged ground-truth, the model is agnostic to the inherent subjectivity in emotional expressions. In this work, we therefore propose an end-to-end Bayesian neural network capable of being trained on a distribution of annotations to also capture the subjectivity-based label uncertainty. Instead of a Gaussian, we model the annotation distribution using Student's$t$-distribution, which also accounts for the number of annotations available. We derive the corresponding Kullback-Leibler divergence loss and use it to train an estimator for the annotation distribution, from which the mean and uncertainty can be inferred. We validate the proposed method using two in-the-wild datasets. We show that the proposed$t$-distribution based approach achieves state-of-the-art uncertainty modeling results in speech emotion recognition, and also consistent results in cross-corpora evaluations. Furthermore, analyses reveal that the advantage of a$t$-distribution over a Gaussian grows with increasing inter-annotator correlation and a decreasing number of annotations available. Navin Raj Prabhu, Nale Lehmann-Willenbrock, Timo Gerkmann |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Multi-Channel Speech Separation Using Spatially Selective Deep Non-Linear FiltersabstractIn a multi-channel separation task with multiple speakers, we aim to recover all individual speech signals from the mixture. In contrast to single-channel approaches, which rely on the different spectro-temporal characteristics of the speech signals, multi-channel approaches should additionally utilize the different spatial locations of the sources for a more powerful separation especially when the number of sources increases. To enhance the spatial processing in a multi-channel source separation scenario, in this work, we propose a deep neural network (DNN) based spatially selective filter (SSF) that can be spatially steered to extract the speaker of interest by initializing a recurrent neural network layer with the target direction. We compare the proposed SSF with a common end-to-end direct separation (DS) approach trained using utterance-wise permutation invariant training (PIT), which only implicitly learns to perform spatial filtering. We show that the SSF has a clear advantage over a DS approach with the same underlying network architecture when there are more than two speakers in the mixture, which can be attributed to a better use of the spatial information. Furthermore, we find that the SSF generalizes much better to additional noise sources that were not seen during training and to scenarios with speakers positioned at a similar angle. Kristina Tesch, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | DriftRec: Adapting Diffusion Models to Blind JPEG RestorationabstractIn this work, we utilize the high-fidelity generation abilities of diffusion models to solve blind JPEG restoration at high compression levels. We propose an elegant modification of the forward stochastic differential equation of diffusion models to adapt them to this restoration task and name our methodDriftRec. Comparing DriftRec against anL2regression baseline with the same network architecture and state-of-the-art techniques for JPEG restoration, we show that our approach can escape the tendency of other methods to generate blurry images, and recovers the distribution of clean images significantly more faithfully. For this, only a dataset of clean/corrupted image pairs and no knowledge about the corruption operation is required, enabling wider applicability to other restoration tasks. In contrast to other conditional and unconditional diffusion models, we utilize the idea that the distributions of clean and corrupted images are much closer to each other than each is to the usual Gaussian prior of the reverse process in diffusion models. Our approach therefore requires only low levels of added noise and needs comparatively few sampling steps even without further optimizations. We show that DriftRec naturally generalizes to realistic and difficult scenarios such as unaligned double JPEG compression and blind restoration of JPEGs found online, without having encountered such examples during training. Simon Welker, Henry N. Chapman, Timo Gerkmann |
IEEE Trans. Image Process. | 3 |
| 2023 | Uncertainty Estimation in Deep Speech Enhancement Using Complex Gaussian Mixture ModelsabstractSingle-channel deep speech enhancement approaches often estimate a single multiplicative mask to extract clean speech without a measure of its accuracy. Instead, in this work, we propose to quantify the uncertainty associated with clean speech estimates in neural network-based speech enhancement. Predictive uncertainty is typically categorized into aleatoric uncertainty and epistemic uncertainty. The former accounts for the inherent uncertainty in data and the latter corresponds to the model uncertainty. Aiming for robust clean speech estimation and efficient predictive uncertainty quantification, we propose to integrate statistical complex Gaussian mixture models (CG-MMs) into a deep speech enhancement framework. More specifically, we model the dependency between input and output stochastically by means of a conditional probability density and train a neural network to map the noisy input to the full posterior distribution of clean speech, modeled as a mixture of multiple complex Gaussian components. Experimental results on different datasets show that the proposed algorithm effectively captures predictive uncertainty and that combining powerful statistical models and deep learning also delivers a superior speech enhancement performance. Huajian Fang, Timo Gerkmann |
ICASSP | 2 |
| 2023 | Partially Adaptive Multichannel Joint Reduction of Ego-Noise and Environmental NoiseabstractHuman-robot interaction relies on a noise-robust audio processing module capable of estimating target speech from audio recordings impacted by environmental noise, as well as self-induced noise, so-called ego-noise. While external ambient noise sources vary from environment to environment, ego-noise is mainly caused by the internal motors and joints of a robot. Egonoise and environmental noise reduction are often decoupled, i.e., ego-noise reduction is performed without considering environmental noise. Recently, a variational autoencoder (VAE)-based speech model has been combined with a fully adaptive non-negative matrix factorization (NMF) noise model to recover clean speech under different environmental noise disturbances. However, its enhancement performance is limited in adverse acoustic scenarios involving, e.g. ego-noise. In this paper, we propose a multichannel partially adaptive scheme to jointly model ego-noise and environmental noise utilizing the VAE-NMF framework, where we take advantage of spatially and spectrally structured characteristics of ego-noise by pre-training the ego-noise model, while retaining the ability to adapt to unknown environmental noise. Experimental results show that our proposed approach outperforms the methods based on a completely fixed scheme and a fully adaptive scheme when ego-noise and environmental noise are present simultaneously. Huajian Fang, Niklas Wittmer, Johannes Twiefel, Stefan Wermter, Timo Gerkmann |
ICASSP | 5 |
| 2023 | Analysing Diffusion-based Generative Approaches Versus Discriminative Approaches for Speech RestorationabstractDiffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks like speech enhancement and dereverberation. While discriminative models have traditionally been argued to be more powerful e.g. for speech enhancement, generative diffusion approaches have recently been shown to narrow this performance gap considerably. In this paper, we systematically compare the performance of generative diffusion models and discriminative approaches on different speech restoration tasks. For this, we extend our prior contributions on diffusion-based speech enhancement in the complex time-frequency domain to the task of bandwith extension. We then compare it to a discriminatively trained neural network with the same network architecture on three restoration tasks, namely speech denoising, dereverberation and bandwidth extension. We observe that the generative approach performs globally better than its discriminative counterpart on all tasks, with the strongest benefit for non-additive distortion models, like in dereverberation and bandwidth extension. Code and audio examples can be found online1. Jean-Marie Lemercier, Julius Richter, Simon Welker, Timo Gerkmann |
ICASSP | 4 |
| 2023 | DiffPhase: Generative Diffusion-Based STFT Phase RetrievalabstractDiffusion probabilistic models have been recently used in a variety of tasks, including speech enhancement and synthesis. As a generative approach, diffusion models have been shown to be especially suitable for imputation problems, where missing data is generated based on existing data. Phase retrieval is inherently an imputation problem, where phase information has to be generated based on the given magnitude. In this work we build upon previous work in the speech domain, adapting a speech enhancement diffusion model specifically for STFT phase retrieval. Evaluation using speech quality and intelligibility metrics shows the diffusion approach is well-suited to the phase retrieval task, with performance surpassing both classical and modern methods. Tal Peer, Simon Welker, Timo Gerkmann |
ICASSP | 3 |
| 2023 | Speech Signal Improvement Using Causal Generative Diffusion ModelsabstractIn this paper, we present a causal speech signal improvement system that is designed to handle different types of distortions. The method is based on a generative diffusion model which has been shown to work well in scenarios with missing data and non-linear corruptions. To guarantee causal processing, we modify the network architecture of our previous work and replace global normalization with causal adaptive gain control. We generate diverse training data containing a broad range of distortions. This work was performed in the context of an "ICASSP Signal Processing Grand Challenge" and submitted to the non-real-time track of the "Speech Signal Improvement Challenge 2023", where it was ranked fifth. Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Tal Peer, Timo Gerkmann |
ICASSP | 6 |
| 2023 | Spatially Selective Deep Non-Linear Filters For Speaker ExtractionabstractIn a scenario with multiple persons talking simultaneously, the spatial characteristics of the signals are the most distinct feature for extracting the target signal. In this work, we develop a deep joint spatial-spectral non-linear filter that can be steered to an arbitrary target direction. For this we propose a simple and effective conditioning mechanism, which sets the initial state of the filter’s recurrent layers based on the target direction. We show that this scheme is more effective than the baseline approach and increases the flexibility of the filter at no performance cost. The resulting spatially selective non-linear filters can also be used for speech separation of an arbitrary number of speakers and enable very accurate multi-speaker localization as we demonstrate in this paper. Kristina Tesch, Timo Gerkmann |
ICASSP | 2 |
| 2023 | Reducing the Prior Mismatch of Stochastic Differential Equations for Diffusion-based Speech Enhancement
Bunlong Lay, Simon Welker, Julius Richter, Timo Gerkmann |
INTERSPEECH | 4 |
| 2023 | Extending DNN-based Multiplicative Masking to Deep Subband Filtering for Improved Dereverberation
Jean-Marie Lemercier, Julian Tobergte, Timo Gerkmann |
INTERSPEECH | 3 |
| 2023 | Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model
Héctor Martel, Julius Richter, Kai Li 0047, Xiaolin Hu 0001, Timo Gerkmann |
INTERSPEECH | 5 |
| 2023 | Leveraging Semantic Information for Efficient Self-Supervised Emotion Recognition with Audio-Textual Distilled ModelsabstractIn large part due to their implicit semantic modeling, selfsupervised learning (SSL) methods have significantly increased the performance of valence recognition in speech emotion recognition (SER) systems.Yet, their large size may often hinder practical implementations.In this work, we take HuBERT as an example of an SSL model and analyze the relevance of each of its layers for SER.We show that shallow layers are more important for arousal recognition while deeper layers are more important for valence.This observation motivates the importance of additional textual information for accurate valence recognition, as the distilled framework lacks the depth of its large-scale SSL teacher.Thus, we propose an audio-textual distilled SSL framework that, while having only ∼20% of the trainable parameters of a large SSL model, achieves on par performance across the three emotion dimensions (arousal, valence, dominance) on the MSP-Podcast v1.10 dataset. Danilo de Oliveira, Navin Raj Prabhu, Timo Gerkmann |
INTERSPEECH | 3 |
| 2023 | Integrating Uncertainty Into Neural Network-Based Speech EnhancementabstractSupervised masking approaches in the time-frequency domain aim to employ deep neural networks to estimate a multiplicative mask to extract clean speech. This leads to a single estimate for each input without any guarantees or measures of reliability. In this paper, we study the benefits of modeling uncertainty in clean speech estimation. Prediction uncertainty is typically categorized intoaleatoric uncertaintyandepistemic uncertainty. The former refers to inherent randomness in data, while the latter describes uncertainty in the model parameters. In this work, we propose a framework to jointly model aleatoric and epistemic uncertainties in neural network-based speech enhancement. The proposed approach captures aleatoric uncertainty by estimating the statistical moments of the speech posterior distribution and explicitly incorporates the uncertainty estimate to further improve clean speech estimation. For epistemic uncertainty, we investigate two Bayesian deep learning approaches: Monte Carlo dropout and Deep ensembles to quantify the uncertainty of the neural network parameters. Our analyses show that the proposed framework promotes capturing practical and reliable uncertainty, while combining different sources of uncertainties yields more reliable predictive uncertainty estimates. Furthermore, we demonstrate the benefits of modeling uncertainty on speech enhancement performance by evaluating the framework on different datasets, exhibiting notable improvement over comparable models that fail to account for uncertainty. Huajian Fang, Dennis Becker, Stefan Wermter, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech Enhancement and DereverberationabstractDiffusion models have shown a great ability at bridging the performance gap between predictive and generative approaches for speech enhancement. We have shown that they may even outperform their predictive counterparts for non-additive corruption types or when they are evaluated on mismatched conditions. However, diffusion models suffer from a high computational burden, mainly as they require to run a neural network for each reverse diffusion step, whereas predictive approaches only require one pass. As diffusion models are generative approaches they may also produce vocalizing and breathing artifacts in adverse conditions. In comparison, in such difficult scenarios, predictive models typically do not produce such artifacts but tend to distort the target speech instead, thereby degrading the speech quality. In this work, we present a stochastic regeneration approach where an estimate given by a predictive model is provided as a guide for further diffusion. We show that the proposed approach uses the predictive model to remove the vocalizing and breathing artifacts while producing very high quality samples thanks to the diffusion model, even in adverse conditions. We further show that this approach enables to use lighter sampling schemes with fewer diffusion steps without sacrificing quality, thus lifting the computational burden by an order of magnitude. Source code and audio examples are available onlinehttps://uhh.de/inf-sp-storm. Jean-Marie Lemercier, Julius Richter, Simon Welker, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Speech Enhancement and Dereverberation With Diffusion-Based Generative ModelsabstractIn this work, we build upon our previous publication and use diffusion-based generative models for speech enhancement. We present a detailed overview of the diffusion process that is based on a stochastic differential equation and delve into an extensive theoretical examination of its implications. Opposed to usual conditional generation tasks, we do not start the reverse process from pure Gaussian noise but from a mixture of noisy speech and Gaussian noise. This matches our forward process which moves from clean speech to noisy speech by including a drift term. We show that this procedure enables using only 30 diffusion steps to generate high-quality clean speech estimates. By adapting the network architecture, we are able to significantly improve the speech enhancement performance, indicating that the network, rather than the formalism, was the main limitation of our original approach. In an extensive cross-dataset evaluation, we show that the improved method can compete with recent discriminative models and achieves better generalization when evaluating on a different corpus than used for training. We complement the results with an instrumental evaluation using real-world noisy recordings and a listening experiment, in which our proposed method is rated best. Examining different sampler configurations for solving the reverse process allows us to balance the performance and computational speed of the proposed method. Moreover, we show that the proposed method is also suitable for dereverberation and thus not limited to additive background noise removal. Code and audio examples are available online1https://github.com/sp-uhh/sgmse. Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Insights Into Deep Non-Linear Filters for Improved Multi-Channel Speech EnhancementabstractThe key advantage of using multiple microphones for speech enhancement is that spatial filtering can be used to complement the tempo-spectral processing. In a traditional setting, linear spatial filtering (beamforming) and single-channel post-filtering are commonly performed separately. In contrast, there is a trend towards employing (DNNs) to learn a joint spatial and tempo-spectral non-linear filter, which means that the restriction of a linear processing model and that of a separate processing of spatial and tempo-spectral information can potentially be overcome. However, the internal mechanisms that lead to good performance of such data-driven filters for multi-channel speech enhancement are not well understood. Therefore, in this work, we analyse the properties of a non-linear spatial filter realized by a DNN as well as its interdependency with temporal and spectral processing by carefully controlling the information sources (spatial, spectral, and temporal) available to the network. We confirm the superiority of a non-linear spatial processing model, which outperforms an oracle linear spatial filter in a challenging speaker extraction scenario for a low number of microphones by 0.24 POLQA score. Our analyses reveal that in particular spectral information should be processed jointly with spatial information as this increases the spatial selectivity of the filter. Our systematic evaluation then leads to a simple network architecture, that outperforms state-of-the-art network architectures on a speaker extraction task by 0.22 POLQA score and by 0.32 POLQA score on the CHiME3 data. Kristina Tesch, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Label Uncertainty Modeling and Prediction for Speech Emotion Recognition using t-DistributionsabstractAs different people perceive others' emotional expressions differently, their annotation in terms of arousal and valence are per se subjective. To address this, these emotion annotations are typically collected by multiple annotators and averaged across annotators in order to obtain labels for arousal and valence. However, besides the average, also the uncertainty of a label is of interest, and should also be modeled and predicted for automatic emotion recognition. In the literature, for simplicity, label uncertainty modeling is commonly approached with a Gaussian assumption on the collected annotations. However, as the number of annotators is typically rather small due to resource constraints, we argue that the Gaussian approach is a rather crude assumption. In contrast, in this work we propose to model the label distribution using a Student's t-distribution which allows us to account for the number of annotations available. With this model, we derive the corresponding Kullback-Leibler divergence based loss function and use it to train an estimator for the distribution of emotion labels, from which the mean and uncertainty can be inferred. Through qualitative and quantitative analysis, we show the benefits of the$t$-distribution over a Gaussian distribution. We validate our proposed method on the AVEC'16 dataset. Results reveal that our$t$-distribution based approach improves over the Gaussian approach with state-of-the-art uncertainty modeling results in speech-based emotion recognition, along with an optimal and even faster convergence. Navin Raj Prabhu, Nale Lehmann-Willenbrock, Timo Gerkmann |
ACII | 3 |
| 2022 | Integrating Statistical Uncertainty into Neural Network-Based Speech EnhancementabstractSpeech enhancement in the time-frequency domain is often performed by estimating a multiplicative mask to extract clean speech. However, most neural network-based methods perform point estimation, i.e., their output consists of a single mask. In this paper, we study the benefits of modeling uncertainty in neural network-based speech enhancement. For this, our neural network is trained to map a noisy spectrogram to the Wiener filter and its associated variance, which quantifies uncertainty, based on the maximum a posteriori (MAP) inference of spectral coefficients. By estimating the distribution instead of the point estimate, one can model the uncertainty associated with each estimate. We further propose to use the estimated Wiener filter and its uncertainty to build an approximate MAP (A-MAP) estimator of spectral magnitudes, which in turn is combined with the MAP inference of spectral coefficients to form a hybrid loss function to jointly reinforce the estimation. Experimental results on different datasets show that the proposed method can not only capture the uncertainty associated with the estimated filters, but also yield a higher enhancement performance over comparable models that do not take uncertainty into account. Huajian Fang, Tal Peer, Stefan Wermter, Timo Gerkmann |
ICASSP | 4 |
| 2022 | Customizable End-To-End Optimization Of Online Neural Network-Supported Dereverberation For Hearing DevicesabstractThis work focuses on online dereverberation for hearing devices using the weighted prediction error (WPE) algorithm. WPE filtering requires an estimate of the target speech power spectral density (PSD). Recently deep neural networks (DNNs) have been used for this task. However, these approaches optimize the PSD estimate which only indirectly affects the WPE output, thus potentially resulting in limited dereverberation. In this paper, we propose an end-to-end approach specialized for online processing, that directly optimizes the dereverberated output signal. In addition, we propose to adapt it to the needs of different types of hearing-device users by modifying the optimization target as well as the WPE algorithm characteristics used in training. We show that the proposed end-to-end approach outperforms the traditional and conventional DNN-supported WPEs on a noise-free version of the WHAMR! dataset. Jean-Marie Lemercier, Joachim Thiemann, Raphael Koning, Timo Gerkmann |
ICASSP | 4 |
| 2022 | Deep Iterative Phase Retrieval for PtychographyabstractOne of the most prominent challenges in the field of diffractive imaging is the phase retrieval (PR) problem: In order to reconstruct an object from its diffraction pattern, the inverse Fourier transform must be computed. This is only possible given the full complex-valued diffraction data, i.e. magnitude and phase. However, in diffractive imaging, generally only magnitudes can be directly measured while the phase needs to be estimated. In this work we specifically consider ptychography, a sub-field of diffractive imaging, where objects are reconstructed from multiple overlapping diffraction images. We pro-pose an augmentation of existing iterative phase retrieval algorithms with a neural network designed for refining the result of each iteration. For this purpose we adapt and extend a recently proposed architecture from the speech processing field. Evaluation results show the proposed approach delivers improved convergence rates in terms of both iteration count and algorithm runtime. Simon Welker, Tal Peer, Henry N. Chapman, Timo Gerkmann |
ICASSP | 4 |
| 2022 | Continuous Phoneme Recognition based on Audio-Visual Modality FusionabstractWhile state-of-the-art audio-only phoneme recognition is already at a high standard, the robustness of existing methods still drops in very noisy environments. To mitigate these limitations, visual information can be incorporated into the recognition system, such that the problem is formulated in a multi-modal setting. To this end, we develop a continuous, audio-visual phoneme classifier that takes raw audio waveforms and video frames as input. Both modalities are processed by individual feature extraction models before a fusion model exploits their correlations. Audio features are extracted with a residual neural network, while video features are obtained with a convolutional neural network. Furthermore, we model temporal dependencies with gated recurrent units. For modality fusion, we compare simple concatenation, attention-based methods, as well as squeeze-and-excitation to learn a joint representation. We train our models on the NTCD-TIMIT dataset, using distinct noise types from the QUT dataset for the test. By pre-training the feature extraction models on the individual modalities first, we achieve best performance for the audio-visual model that is trained end-to-end. In the experiments, we show that by including the video modality, we increase the accuracy of phoneme prediction by 9% in very noisy acoustic environments. The results indicate that in such environments our approach remains more robust compared to existing methods. The code and pre-trained models are available online11https://github.com/sp-uhh/av-phoneme. Julius Richter, Jeanine Liebold, Timo Gerkmann |
IJCNN | 3 |
| 2022 | Neural Network-augmented Kalman Filtering for Robust Online Speech Dereverberation in Noisy Reverberant EnvironmentsabstractIn this paper, a neural network-augmented algorithm for noise-robust online dereverberation with a Kalman filtering variant of the weighted prediction error (WPE) method is proposed.The filter stochastic variations are predicted by a deep neural network (DNN) trained end-to-end using the filter residual error and signal characteristics.The presented framework allows for robust dereverberation on a single-channel noisy reverberant dataset similar to WHAMR!.The Kalman filtering WPE introduces distortions in the enhanced signal when predicting the filter variations from the residual error only, if the target speech power spectral density is not perfectly known and the observation is noisy.The proposed approach avoids these distortions by correcting the filter variations estimation in a data-driven way, increasing the robustness of the method to noisy scenarios.Furthermore, it yields a strong dereverberation and denoising performance compared to a DNN-supported recursive least squares variant of WPE, especially for highly noisy inputs. Jean-Marie Lemercier, Joachim Thiemann, Raphael Koning, Timo Gerkmann |
INTERSPEECH | 4 |
| 2022 | Efficient Transformer-based Speech Enhancement Using Long Frames and STFT MagnitudesabstractThe SepFormer architecture shows very good results in speech separation. Like other learned-encoder models, it uses short frames, as they have been shown to obtain better performance in these cases. This results in a large number of frames at the input, which is problematic; since the SepFormer is transformer-based, its computational complexity drastically increases with longer sequences. In this paper, we employ the SepFormer in a speech enhancement task and show that by replacing the learned-encoder features with a magnitude short-time Fourier transform (STFT) representation, we can use long frames without compromising perceptual enhancement performance. We obtained equivalent quality and intelligibility evaluation scores while reducing the number of operations by a factor of approximately 8 for a 10-second utterance. Danilo de Oliveira, Tal Peer, Timo Gerkmann |
INTERSPEECH | 3 |
| 2022 | End-To-End Label Uncertainty Modeling for Speech-based Arousal Recognition Using Bayesian Neural NetworksabstractEmotions are subjective constructs.Recent end-to-end speech emotion recognition systems are typically agnostic to the subjective nature of emotions, despite their state-of-the-art performance.In this work, we introduce an end-to-end Bayesian neural network architecture to capture the inherent subjectivity in the arousal dimension of emotional expressions.To the best of our knowledge, this work is the first to use Bayesian neural networks for speech emotion recognition.At training, the network learns a distribution of weights to capture the inherent uncertainty related to subjective arousal annotations.To this end, we introduce a loss term that enables the model to be explicitly trained on a distribution of annotations, rather than training them exclusively on mean or gold-standard labels.We evaluate the proposed approach on the AVEC'16 dataset.Qualitative and quantitative analysis of the results reveals that the proposed model can aptly capture the distribution of subjective arousal annotations, with state-of-the-art results in mean and standard deviation estimations for uncertainty modeling. Navin Raj Prabhu, Guillaume Carbajal, Nale Lehmann-Willenbrock, Timo Gerkmann |
INTERSPEECH | 4 |
| 2022 | On the Role of Spatial, Spectral, and Temporal Processing for DNN-based Non-linear Multi-channel Speech Enhancement
Kristina Tesch, Nils-Hendrik Mohrmann, Timo Gerkmann |
INTERSPEECH | 3 |
| 2022 | Speech Enhancement with Score-Based Generative Models in the Complex STFT DomainabstractScore-based generative models (SGMs) have recently shown impressive results for difficult generative tasks such as the unconditional and conditional generation of natural images and audio signals.In this work, we extend these models to the complex short-time Fourier transform (STFT) domain, proposing a novel training task for speech enhancement using a complex-valued deep neural network.We derive this training task within the formalism of stochastic differential equations (SDEs), thereby enabling the use of predictor-corrector samplers.We provide alternative formulations inspired by previous publications on using generative diffusion models for speech enhancement, avoiding the need for any prior assumptions on the noise distribution and making the training task purely generative which, as we show, results in improved enhancement performance. Simon Welker, Julius Richter, Timo Gerkmann |
INTERSPEECH | 3 |
| 2022 | Speech Enhancement Regularized by a Speaker Verification ModelabstractIn the last years, end-to-end neural networks were employed to enhance single-channel noisy speech. The Conv-Tasnet architecture is such a neural network and it has been successfully trained on the scale invariant signal-to-distortion ratio (SI-SDR) loss. However, we find that in a speech enhancement task Conv-Tasnet trained on the SI-SDR loss introduces distortions to the enhanced signal as the mismatch between training and testing data increases. To mitigate these effects, we investigate on a new loss function, that combines SI-SDR with a pretrained speaker verification (SV) model as a regularizer. As the SV model captures information about important characteristics of clean speech signals, we argue that the proposed regularization ensures that the estimated speech signal is speech-like even if there is an increasing mismatch between training and testing data. Accordingly, we observe considerable improvements in POLQA in mismatched training and testing conditions, e.g. when training on anechoic speech but testing on reverberant data. Bunlong Lay, Timo Gerkmann |
MMSP | 2 |
| 2021 | Guided Variational Autoencoder for Speech Enhancement with a Supervised ClassifierabstractRecently, variational autoencoders have been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. However, variational autoencoders are trained on clean speech only, which results in a limited ability of extracting the speech signal from noisy speech compared to supervised approaches. In this paper, we propose to guide the variational autoencoder with a supervised classifier separately trained on noisy speech. The estimated label is a high-level categorical variable describing the speech signal (e.g. speech activity) allowing for a more informed latent distribution compared to the standard variational autoencoder. We evaluate our method with different types of labels on real recordings of different noisy environments. Provided that the label better informs the latent distribution and that the classifier achieves good performance, the proposed approach outperforms the standard variational autoencoder and a conventional neural network- based supervised approach. Guillaume Carbajal, Julius Richter, Timo Gerkmann |
ICASSP | 3 |
| 2021 | Variational Autoencoder for Speech Enhancement with a Noise-Aware EncoderabstractRecently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to noise presence, especially in low signal-to-noise ratios (SNRs). To increase the robustness of the VAE, we propose to include noise information in the training phase by using a noise-aware encoder trained on noisy-clean speech pairs. We evaluate our approach on real recordings of different noisy environments and acoustic conditions using two different noise datasets. We show that our proposed noise-aware VAE outperforms the standard VAE in terms of overall distortion without increasing the number of model parameters. At the same time, we demonstrate that our model is capable of generalizing to unseen noise conditions better than a supervised feedforward deep neural network (DNN). Furthermore, we demonstrate the robustness of the model performance to a reduction of the noisy-clean speech training data size. Huajian Fang, Guillaume Carbajal, Stefan Wermter, Timo Gerkmann |
ICASSP | 4 |
| 2021 | See the Silence: Improving Visual-Only Voice Activity Detection by Optical Flow and RGB Fusion
Danu Caus, Guillaume Carbajal, Timo Gerkmann, Simone Frintrop |
ICVS | 3 |
| 2021 | Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkabstractRecent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired architecture called Fully Recurrent Convolutional Neural Network (FRCNN) to solve the separation task. This model contains bottom-up, top-down and lateral connections to fuse information processed at various time-scales represented by stages. In contrast to the traditional approach updating stages in parallel, we propose to first update the stages one by one in the bottom-up direction, then fuse information from adjacent stages simultaneously and finally fuse information from all stages to the bottom stage together. Experiments showed that this asynchronous updating scheme achieved significantly better results with much fewer parameters than the traditional synchronous updating scheme on speech separation. In addition, the proposed model achieved competitive or better results with high efficiency as compared to other state-of-the-art approaches on two benchmark datasets. Xiaolin Hu 0001, Kai Li 0047, Jean-Marie Lemercier, Timo Gerkmann |
NeurIPS | 6 |
| 2021 | SNR-Based Features and Diverse Training Data for Robust DNN-Based Speech EnhancementabstractIn this paper, we address the generalization of deep neural network (DNN) based speech enhancement to unseen noise conditions for the case that training data is limited in size and diversity. To gain more insights, we analyze the generalization with respect to (1) the size and diversity of the training data, (2) different network architectures, and (3) the chosen features. To address (1), we train networks on the Hu noise corpus (limited size), the CHiME 3 noise corpus (limited diversity) and also propose a large and diverse dataset collected based on freely available sounds. To address (2), we compare a fully-connected feed-forward and a long short-term memory (LSTM) architecture. To address (3), we compare three input features, namely logarithmized noisy periodograms, noise aware training (NAT) and the proposed signal-to-noise ratio based noise aware training (SNR-NAT). We confirm that rich training data and improved network architectures help DNNs to generalize. Furthermore, we show via experimental results and an analysis using t-distributed stochastic neighbor embedding (t-SNE) that the proposed SNR-NAT features yield robust and level independent results in unseen noise even with simple network architectures and when trained on only small datasets, which is the key contribution of this paper. Robert Rehr, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Nonlinear Spatial Filtering in Multichannel Speech EnhancementabstractThe majority of multichannel speech enhancement algorithms are two-step procedures that first apply a linear spatial filter, a so-called beamformer, and combine it with a single-channel approach for postprocessing. However, the serial concatenation of a linear spatial filter and a postfilter is not generally optimal in the minimum mean square error (MMSE) sense for noise distributions other than a Gaussian distribution. Rather, the MMSE optimal filter is ajoint spatial and spectral nonlinearfunction. While estimating the parameters of such a filter with traditional methods is challenging, modern neural networks may provide an efficient way to learn the nonlinear function directly from data. To see if further research in this direction is worthwhile, in this work we examine the potential performance benefit of replacing the common two-step procedure with a joint spatial and spectral nonlinear filter. We analyze three different forms of non-Gaussianity: First, we evaluate on super-Gaussian noise with a high kurtosis. Second, we evaluate on inhomogeneous noise fields created by five interfering sources using two microphones, and third, we evaluate on real-world recordings from the CHiME3 database. In all scenarios, considerable improvements may be obtained. Most prominently, our analyses show that a nonlinear spatial filter uses the available spatial information more effectively than a linear spatial filter as it is capable of suppressing more than$D-1$directional interfering sources with a$D$-dimensional microphone array without spatial adaptation. Kristina Tesch, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Efficient Joint Estimation of Tracer Distribution and Background Signals in Magnetic Particle Imaging Using a Dictionary ApproachabstractBackground signals are a primary source of artifacts in magnetic particle imaging and limit the sensitivity of the method since background signals are often not precisely known and vary over time. The state-of-the art method for handling background signals uses one or several background calibration measurements with an empty scanner bore and subtracts a linear combination of these background measurements from the actual particle measurement. This approach yields satisfying results in case that the background measurements are taken in close proximity to the particle measurement and when the background signal drifts linearly. In this work, we propose a joint estimation of particle distribution and background signal based on a dictionary that is capable of representing typical background signals. Reconstruction is performed frame-by-frame with minimal assumptions on the temporal evolution of background signals. Thus, even non-linear temporal evolution of the latter can be captured. Using a singular-value decomposition, the dictionary is derived from a large number of background calibration scans that do not need to be recorded in close proximity to the particle measurement. The dictionary is sufficiently expressive and represented by its principle components. The proposed joint estimation of particle distribution and background signal is expressed as a linear Tikhonov-regularized least squares problem, which can be efficiently solved. In phantom experiments it is shown that the method strongly suppresses background artifacts and even allows to estimate and remove the direct feed-through of the excitation field. Tobias Knopp 0001, Mirco Grosser, Matthias Graeser, Timo Gerkmann, Martin Möddel |
IEEE Trans. Medical Imaging | 4 |
| 2020 | A Multi-Phase Gammatone Filterbank for Speech Separation Via TasnetabstractIn this work, we investigate if the learned encoder of the end-to-end convolutional time domain audio separation network (Conv-TasNet) is the key to its recent success, or if the encoder can just as well be replaced by a deterministic hand-crafted filterbank. Motivated by the resemblance of the trained encoder of Conv-TasNet to auditory filterbanks, we propose to employ a deterministic gammatone filterbank. In contrast to a common gammatone filterbank, our filters are restricted to 2 ms length to allow for low-latency processing. Inspired by the encoder learned by Conv-TasNet, in addition to the logarithmically spaced filters, the proposed filterbank holds multiple gammatone filters at the same center frequency with varying phase shifts. We show that replacing the learned encoder with our proposed multi-phase gammatone filterbank (MP-GTF) even leads to a scale-invariant source-to-noise ratio (SI-SNR) improvement of 0.7 dB. Furthermore, in contrast to using the learned encoder we show that the number of filters can be reduced from 512 to 128 without loss of performance. David Ditter, Timo Gerkmann |
ICASSP | 2 |
| 2020 | Nonlinear Spatial Filtering for Multichannel Speech Enhancement in Inhomogeneous Noise FieldsabstractA common processing pipeline for multichannel speech enhancement is to combine a linear spatial filter with a single-channel postfilter. In fact, it can be shown that such a combination is optimal in the minimum mean square error (MMSE) sense if the noise follows a multivariate Gaussian distribution. However, for non-Gaussian noise, this serial concatenation is generally suboptimal and may thus also lead to suboptimal results. For instance, in our previous work, we showed that a joint spatial-spectral nonlinear estimator achieves a performance gain of 2.6 dB segmental signal-to-noise ratio (SNR) improvement for heavy-tailed large-kurtosis multivariate noise compared to the traditional combination of a linear spatial beamformer and a postfilter.In this paper, we show that a joint spatial-spectral nonlinear filter is not only advantageous for noise distributions that are significantly more heavy-tailed than a Gaussian but also for distributions that model inhomogeneous noise fields while having rather low kurtosis. In experiments with artificially created noise we measure a gain of 1 dB for inhomogenous noise with low kurtosis and up to 2 dB for inhomogeneous noise fields with moderate kurtosis. Kristina Tesch, Timo Gerkmann |
ICASSP | 2 |
| 2020 | Improving mix-and-separate training in audio-visual sound source separation with an object priorabstractThe performance of an audio-visual sound source separation system is determined by its ability to separate audio sources given the images of the sources and the audio mixture. The goal of this study is to investigate the ability to learn the mapping between the sounds and the images of instruments in the self-supervisied mix-and-seperate training paradigm used by state-of-the-art audio-visual sound source separation methods. Theoretical and empirical analyses illustrate that the self-supervised mix-and-separate training does not automatically learn the 1-to-1 correspondence between visual and audio signals, leading to low audio-visual object classification accuracy. Based on this analysis, a weakly-supervised method called Object-Prior is proposed and evaluated on two audio-visual datasets. The experimental results show that the Object-Prior method outperforms state-of-the-art baselines in the audio-visual sound source separation task. It is also more robust against asynchronized data, where the frame and the audio do not come from the same video, and recognizes musical instruments based on their sound with higher accuracy. This indicates that learning the 1-to-1 correspondence between visual and audio features of an instrument improves the effectiveness of audio-visual sound source separation. Julius Richter, Mikko Lauri, Timo Gerkmann, Simone Frintrop |
ICPR | 4 |
| 2020 | Speech Enhancement with Stochastic Temporal Convolutional Networks
Julius Richter, Guillaume Carbajal, Timo Gerkmann |
INTERSPEECH | 3 |
| 2020 | Robust Robotic Pouring using Audition and HapticsabstractRobust and accurate estimation of liquid height lies as an essential part of pouring tasks for service robots. However, vision-based methods often fail in occluded conditions while audio-based methods cannot work well in a noisy environment. We instead propose a multimodal pouring network (MP-Net) that is able to robustly predict liquid height by conditioning on both audition and haptics input. MP-Net is trained on a self-collected multimodal pouring dataset. This dataset contains 300 robot pouring recordings with audio and force/torque measurements for three types of target containers. We also augment the audio data by inserting robot noise. We evaluated MP-Net on our collected dataset and a wide variety of robot experiments. Both network training results and robot experiments demonstrate that MP-Net is robust against noise and changes to the task and environment. Moreover, we further combine the predicted height and force data to estimate the shape of the target container. Hongzhuo Liang, Chuangchuang Zhou, Shuang Li 0014, Xiaojian Ma 0001, Norman Hendrich, Timo Gerkmann, Fuchun Sun 0001, Marcus Stoffel, Jianwei Zhang 0001 |
IROS | 6 |
| 2019 | An Analysis of Noise-aware Features in Combination with the Size and Diversity of Training Data for DNN-based Speech EnhancementabstractIn this work, the generalization of speech enhancement algorithms based on deep neural networks (DNNs) for training datasets that differ in size and diversity is analyzed. For this, we compare noise aware training (NAT) features and signal-to-noise ratio (SNR) based noise aware training (SNR-NAT) features. NAT appends an estimate of the noise power spectral density (PSD) to a noisy periodogram input feature, whereas SNR-NAT uses the noise PSD for normalization. We show that the Hu noise corpus (limited size) and the CHiME 3 noise corpus (limited diversity) may result in DNNs which do not generalize well to unseen noises. We construct a large and diverse dataset from freely available data and show that it helps DNNs to generalize. However, we also show that with SNR-NAT features, the trained models are more robust even if a small or less diverse training set is employed. Using t-distributed stochastic neighbor embedding (t-SNE), we demonstrate that using SNR-NAT both the features and the resulting internal representation of the DNN are less dependent on the background noise which facilitates the generalization to unseen noise types. Robert Rehr, Timo Gerkmann |
ICASSP | 2 |
| 2019 | Influence of Speaker-Specific Parameters on Speech Separation Systems
David Ditter, Timo Gerkmann |
INTERSPEECH | 2 |
| 2019 | On Nonlinear Spatial Filtering in Multichannel Speech EnhancementabstractThe majority of multichannel speech enhancement algorithms are two-step procedures that first apply a linear spatial filter, a so-called beamformer, and combine it with a single-channel approach for postprocessing.However, the serial concatenation of a linear spatial filter and a postfilter is not generally optimal in the minimum mean square error (MMSE) sense for noise distributions other than a Gaussian distribution.Rather, the MMSE optimal filter is a joint spatial and spectral nonlinear function.While estimating the parameters of such a filter with traditional methods is challenging, modern neural networks may provide an efficient way to learn the nonlinear function directly from data.To see if further research in this direction is worthwhile, in this work we examine the potential performance benefit of replacing the common two-step procedure with a joint spatial and spectral nonlinear filter.We analyze three different forms of non-Gaussianity: First, we evaluate on super-Gaussian noise with a high kurtosis.Second, we evaluate on inhomogeneous noise fields created by five interfering sources using two microphones, and third, we evaluate on realworld recordings from the CHiME3 database.In all scenarios, considerable improvements may be obtained.Most prominently, our analyses show that a nonlinear spatial filter uses the available spatial information more effectively than a linear spatial filter as it is capable of suppressing more than D -1 directional interfering sources with a D-dimensional microphone array without spatial adaptation. Kristina Tesch, Robert Rehr, Timo Gerkmann |
INTERSPEECH | 3 |
| 2019 | Making Sense of Audio Vibration for Liquid Height Estimation in Robotic PouringabstractIn this paper, we focus on the challenging perception problem in robotic pouring. Most of the existing approaches either leverage visual or haptic information. However, these techniques may suffer from poor generalization performances on opaque containers or concerning measuring precision. To tackle these drawbacks, we propose to make use of audio vibration sensing and design a deep neural network PouringNet to predict the liquid height from the audio fragment during the robotic pouring task. PouringNet is trained on our collected real-world pouring dataset with multimodal sensing data, which contains more than 3000 recordings of audio, force feedback, video and trajectory data of the human hand that performs the pouring task. Each record represents a complete pouring procedure. We conduct several evaluations on PouringNet with our dataset and robotic hardware. The results demonstrate that our PouringNet generalizes well across different liquid containers, positions of the audio receiver, initial liquid heights and types of liquid, and facilitates a more robust and accurate audio-based perception for robotic pouring. Hongzhuo Liang, Shuang Li 0014, Xiaojian Ma 0001, Norman Hendrich, Timo Gerkmann, Fuchun Sun 0001, Jianwei Zhang 0001 |
IROS | 5 |
| 2018 | Nonlinear Speech Enhancement Under Speech PSD UncertaintyabstractMost Bayesian clean speech estimators, like the Wiener filter or Ephraim and Malah's amplitude estimators, are derived under the assumption that the true power spectral density (PSD) of speech is known. In practice, however, only estimates are available. When the PSD estimation errors are neglected, they propagate through to the final speech estimate, resulting in undesired artifacts such as musical noise and speech distortions. To increase the robustness to PSD estimation errors, recently a linear estimator has been proposed that explicitly takes into account the uncertainty of the available speech PSD estimate. In this paper, we show that in the derivation of this estimator a limiting statistical assumption is made, and that avoiding this assumption leads to a novel, potentially more powerful nonlinear estimator under PSD uncertainty. In combination with a sophisticated speech PSD estimator, the proposed approach achieves a higher predicted speech quality than the linear alternative and its conventional counterpart, the Wiener filter. Martin Krawczyk-Becker, Timo Gerkmann |
ICASSP | 2 |
| 2018 | Weighted and Multi-Task Loss for Rare Audio Event DetectionabstractWe present in this paper two loss functions tailored for rare audio event detection in audio streams. The weighted loss is designed to tackle the common issue of imbalanced data in background/foreground classification while the multi-task loss enables the networks to simultaneously model the class distribution and the temporal structures of the target events for recognition. We study the proposed loss functions with deep neural networks (DNNs) and convolutional neural networks (CNNs) coupled with state-of-the-art phase-aware signal enhancement. Experiments on the DCASE 2017 challenge's data show that our system with the proposed losses significantly outperforms not only the DCASE 2017 baseline but also our baseline which has a similar network architecture and a standard loss function. Huy Phan, Martin Krawczyk-Becker, Timo Gerkmann, Alfred Mertins |
ICASSP | 3 |
| 2018 | On Speech Enhancement Under PSD UncertaintyabstractMany well-known and frequently employed Bayesian clean speech estimators have been derived under the assumption that the true power spectral densities (PSDs) of speech and noise are exactly known. In practice, however, only power spectral density (PSD) estimates are available. Simply neglecting PSD estimation errors and handling the estimates as true values leads to speech estimation errors causing musical noise and undesired suppression of speech. In this paper, the uncertainty of the available speech PSD estimates is addressed. The main contributions are the following. First, we summarize and examine ways to model and incorporate the uncertainty of PSD estimates for a more robust speech enhancement performance. Second, a novel nonlinear clean speech estimator is derived that takes into account prior knowledge about the absolute value of typical speech PSDs. Third, we show that the derived statistical framework provides uncertainty-aware counterparts to a number of well-known conventional clean speech estimators such as the Wiener filter and Ephraim and Malah's amplitude estimators. Fourth, we show how modern PSD estimators can be incorporated into the theoretical framework and propose to employ frequency dependent priors. Finally, the effects and benefits of considering the uncertainty of speech PSD estimates are analyzed, discussed, and evaluated via instrumental measures and a listening experiment. Martin Krawczyk-Becker, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | On the Importance of Super-Gaussian Speech Priors for Machine-Learning Based Speech EnhancementabstractFor enhancing noisy signals, machine-learning based single-channel speech enhancement schemes exploit prior knowledge about typical speech spectral structures. To ensure a good generalization and to meet requirements in terms of computational complexity and memory consumption, certain methods restrict themselves to learning speech spectral envelopes. We refer to these approaches as machine-learning spectral envelope (MLSE)-based approaches. In this paper, we show by means of theoretical and experimental analyses that for MLSE-based approaches, super-Gaussian priors allow for a reduction of noise between speech spectral harmonics which is not achievable using Gaussian estimators such as the Wiener filter. For the evaluation, we use a deep neural network based phoneme classifier and a low-rank nonnegative matrix factorization framework as examples of MLSE-based approaches. A listening experiment and instrumental measures confirm that while super-Gaussian priors yield only moderate improvements for classic enhancement schemes, for MLSE-based approaches super-Gaussian priors clearly make an important difference and significantly outperform Gaussian priors. Robert Rehr, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | MixMax Approximation as a Super-Gaussian Log-Spectral Amplitude Estimator for Speech Enhancement
Robert Rehr, Timo Gerkmann |
INTERSPEECH | 2 |
| 2017 | An Analysis of Adaptive Recursive Smoothing with Applications to Noise PSD EstimationabstractFirst-order recursive smoothing filters using a fixed smoothing constant are in general unbiased estimators of the mean of a random process. Due to their efficiency in terms of memory consumption and computational complexity, they are of high practical relevance and are also often used to track the first-order moment of nonstationary random processes. However, in single-channel speech-enhancement applications, e.g., for the estimation of the noise power spectral density, an adaptively changing smoothing factor is often employed. Here, the adaptivity is used to avoid speech leakage by raising the smoothing factor when speech is likely to be present. In this paper, we investigate the properties of adaptive first-order recursive smoothing factors applied to noise power spectral density estimators. We show that in contrast to a smoothing with fixed smoothing factors, adaptive smoothing is in general biased. We propose different methods to quantify and to compensate for the bias. We demonstrate that the proposed correction methods reduce the estimation error and increases the perceptual evaluation of speech quality scores in a speech enhancement framework. Robert Rehr, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Sparse reconstruction of quantized speech signalsabstractWe propose sparse reconstruction techniques to improve the quality and/or reduce the bit-rate of standard speech coders. To that end, we assume signal sparsity in some transform domain and formulate the problem of reconstructing the original signal in terms of constrained ℓ1-norm minimization. We use modern primal-dual methods in order to solve the resulting non-smooth convex optimization problem. Experiments show that with the proposed sparse reconstruction method the instrumentally predicted speech quality can be largely improved. Christoph Brauer, Timo Gerkmann, Dirk A. Lorenz |
ICASSP | 2 |
| 2016 | Perceptual and instrumental evaluation of the perceived level of reverberationabstractPerceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significant. Therefore, an efficient perceptual measure of the perceived level of reverberation is needed. We compare the use of a multiple stimuli test with the use of pairwise comparison for the evaluation of the perceived level of reverberation. The results suggest that using multiple stimuli is preferable to pairwise comparison as long as the number of conditions to be compared is not too large. Additionally, we use the results from the conducted perceptual measurements to examine the reliability of existing instrumental measures of the perceived level of reverberation. Our observations show which instrumental measures are effective in highlighting differences between RIR characteristics and which ones have to be preferred if one aims at predicting the level of reverberation perceived by a human assessor. Benjamin Cauchi, Hamza A. Javed, Timo Gerkmann, Simon Doclo, Stefan Goetze, Patrick A. Naylor |
ICASSP | 3 |
| 2016 | Single-microphone speech enhancement using MVDR filtering and Wiener post-filteringabstractFor single-microphone noise reduction, a minimum variance distortionless response (MVDR) filter has been proposed recently. This filter takes the speech correlations of consecutive time frames into account and achieves impressive results in terms of speech distortions even in a blind implementation where we only have access to the noisy speech signal. However, compared to conventional approaches less noise reduction is achieved. Therefore, we propose to combine the single-microphone MVDR with a Wiener post-filter as the minimum-mean-square error optimal solution when multiple time frames are considered. We propose to pre-train the required interframe coherence matrices of the interferences for a large database, while speech correlations and interference power spectral densities are estimated online. In an experimental study based on instrumental measures, the proposed approach achieves a good trade-off between a single-channel Wiener filter and a multi-frame MVDR. Dörte Fischer, Timo Gerkmann |
ICASSP | 2 |
| 2016 | BIAS correction methods for adaptive recursive smoothing with applications in noise PSD estimationabstractDue to the low computational complexity and the low memory consumption, first-order recursive smoothing is a technique often applied to estimate the mean of a random process. For instance, recursive smoothing is used in noise power estimators where adaptively changing smoothing factors are used instead of fixed ones to prevent the speech power from leaking into the noise estimate. However, in general, the usage of adaptive smoothing factors leads to a biased estimate of the mean. In this paper, we propose a novel method to correct the bias evoked by adaptive smoothing factors. We compare this method to a recently proposed compensation method in terms of the log-error distortion using real world signals for two noise power estimators. We show that both corrections reduce the distortion measure in noisy speech while the novel method has the advantage that no iteration is required for determining the correction factor. Robert Rehr, Timo Gerkmann |
ICASSP | 2 |
| 2016 | Fundamental Frequency Informed Speech Enhancement in a Flexible Statistical FrameworkabstractConventional statistical clean speech estimators, like the Wiener filter, are frequently used for the spectro-temporal enhancement of noise corrupted speech. Most of these approaches estimate the clean speech independently for each time-frequency point, neglecting the structure of the underlying speech sound. In this work, we derive a statistical estimator that explicitly takes into account information about the characteristic structure of voiced speech by means of a harmonic signal model. To this end, we also present a way to estimate a harmonic model-based clean speech representation and the corresponding error variance directly in the short-time Fourier transform domain. The resulting estimator is optimal in the minimum-mean-squared error sense and can conveniently be formulated in terms of a multichannel Wiener filter. The proposed estimator outperforms several reference algorithms in terms of speech quality and intelligibility as predicted by instrumental measures. Martin Krawczyk-Becker, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | On MMSE-Based Estimation of Amplitude and Complex Speech Spectral Coefficients Under Phase-UncertaintyabstractAmong the most commonly used single-channel approaches for the enhancement of noise corrupted speech are Bayesian estimators of clean speech coefficients in the short-time Fourier transform domain. However, the vast majority of these approaches effectively only modifies the spectral amplitude and does not consider any information about the clean speech spectral phase. More recently, clean speech estimators that can utilize prior phase information have been proposed and shown to lead to improvements over the traditional, phase-blind approaches. In this work, we revisit phase-aware estimators of clean speech amplitudes and complex coefficients. To complete the existing set of estimators, we first derive a novel amplitude estimator given uncertain prior phase information. Second, we derive a closed-form solution for complex coefficients when the prior phase information is completely uncertain or not available. We put the novel estimators into the context of existing estimators and discuss their advantages and disadvantages. Martin Krawczyk-Becker, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximationabstractIn many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transform domain, which is typically performed in each frequency bin independently without taking into account the spectral structure of the speech signal. Since it is widely accepted that a speech spectrogram can be well approximated with a low-rank matrix, e.g., using a spectral dictionary, in this paper we propose to incorporate a low-rank matrix approximation of the speech spectrogram into the MCLP-based speech dereverberation. The low-rank approximation is obtained using nonnegative matrix factorization with Itakura-Saito divergence. Experimental results for several measured acoustic systems show that incorporating a low-rank approximation improves the dereverberation performance in terms of instrumental speech quality measures. Ante Jukic, Nasser Mohammadiha, Toon van Waterschoot, Timo Gerkmann, Simon Doclo |
ICASSP | 4 |
| 2015 | Utilizing spectro-temporal correlations for an improved speech presence probability based noise power estimationabstractFor the enhancement of speech degraded by noise, accurate estimation of the noise power spectral density (PSD) is indispensable, especially if only a single microphone signal is available. Fast and accurate tracking of the noise PSD is particularly challenging in highly non-stationary noise types, since the distinction between speech and noise components becomes more difficult. Short-time discrete Fourier transform (STFT) based noise PSD estimation algorithms which employ estimates of the speech presence probability (SPP) with fixed priors have been shown to yield good tracking performance even in adverse noise conditions. In this paper, we compare two methods to incorporate spectro-temporal correlations to improve the tracking performance. The first method smoothes the noisy observation over time and frequency before computing the SPP, while the second is based on a Hidden Markov Model (HMM) of the speech presence and absence states. We show that the proposed modifications lead to improved noise PSD estimators which are less sensitive to spectral outliers of the noise and track changes in the noise PSD more quickly than the reference method. Further, when employed in a common speech enhancement setup, the proposed estimators achieve an increased noise reduction while keeping speech distortions at a comparable level. Martin Krawczyk-Becker, Dörte Fischer, Timo Gerkmann |
ICASSP | 3 |
| 2015 | Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparisonabstractIn this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities (PSDs). We first derive closedform expressions for the mean square error (MSE) of both PSD estimators and then show that one estimator - previously used for speech dereverberation by the authors - always yields a better MSE. Only in the case of a two microphone array or for special spatial distributions of the interference both estimators yield the same MSE. The theoretically derived MSE values are in good agreement with numerical simulation results and with instrumental speech quality measures in a realistic speech dereverberation task for binaural hearing aids. Adam Kuklasinski, Simon Doclo, Timo Gerkmann, Søren Holdt Jensen, Jesper Jensen 0001 |
ICASSP | 3 |
| 2015 | Cepstral noise subtraction for robust automatic speech recognitionabstractThe robustness of speech recognizers towards noise can be increased by normalizing the statistical moments of the Mel-frequency cepstral coefficients (MFCCs), e. g. by using cepstral mean normalization (CMN) or cepstral mean and variance normalization (CMVN). The necessary statistics are estimated over a long time window and often, a complete utterance is chosen. Consequently, changes in the background noise can only be tracked to a limited extent which poses a restriction to the performance gain that can be achieved by these techniques. In contrast, algorithms recently developed for single-channel speech enhancement allow to track the background noise quickly. In this paper, we aim at combining speech enhancement techniques and feature normalization methods. For this, we propose to transform an estimate of the noise power spectral density to the MFCC domain, where we subtract it from the noisy MFCCs. This is followed by a conventional CMVN. For background noises that are too instationary for CMVN but can be tracked by the noise estimator, we show that this processing leads to an improvement in comparison to the sole application of CMVN. The observed performance gain emerges especially in low signal-to-noise-ratios. Robert Rehr, Timo Gerkmann |
ICASSP | 2 |
| 2015 | Least squares estimate of the initial phases in STFT based speech enhancementabstractIn this paper, we consider single-channel speech enhancement in the short time Fourier transform (STFT) domain. We suggest to improve an STFT phase estimate by estimating the initial phases. The method is based on the harmonic model and a model for the phase evolution over time. The initial phases are estimated by setting up a least squares problem between the noisy phase and the model for phase evolution. Simulations on synthetic and speech signals show a decreased error on the phase when an estimate of the initial phase is included compared to using the noisy phase as an initialisation. The error on the phase is decreased at input SNRs from -10 to 10 dB. Reconstructing the signal using the clean amplitude, the mean squared error is decreased and the PESQ score is increased. Sidsel Marie Nørholm, Martin Krawczyk-Becker, Timo Gerkmann, Steven van de Par, Jesper Rindom Jensen, Mads Græsbøll Christensen |
INTERSPEECH | 3 |
| 2015 | Multi-Channel Linear Prediction-Based Speech Dereverberation With Sparse PriorsabstractThe quality of speech signals recorded in an enclosure can be severely degraded by room reverberation. In this paper, we focus on a class of blind batch methods for speech dereverberation in a noiseless scenario with a single source, which are based on multi-channel linear prediction in the short-time Fourier transform domain. Dereverberation is performed by maximum-likelihood estimation of the model parameters that are subsequently used to recover the desired speech signal. Contrary to the conventional method, we propose to model the desired speech signal using a general sparse prior that can be represented in a convex form as a maximization over scaled complex Gaussian distributions. The proposed model can be interpreted as a generalization of the commonly used time-varying Gaussian model. Furthermore, we reformulate both the conventional and the proposed method as an optimization problem with an lp-norm cost function, emphasizing the role of sparsity in the considered speech dereverberation methods. Experimental evaluation in different acoustic scenarios show that the proposed approach results in an improved performance compared to the conventional approach in terms of instrumental measures for speech quality. Ante Jukic, Toon van Waterschoot, Timo Gerkmann, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Two-Stage Filter-Bank System for Improved Single-Channel Noise Reduction in Hearing AidsabstractThe filter-bank system implemented in hearing aids has to fulfill various constraints such as low latency and high stop-band attenuation, usually at the cost of low frequency resolution. In the context of frequency-domain noise-reduction algorithms, insufficient frequency resolution may lead to annoying residual noise artifacts since the spectral harmonics of the speech cannot properly be resolved. Especially in case of female speech signals, the noise between the spectral harmonics causes a distinct roughness of the processed signals. Therefore, this work proposes a two-stage filter-bank system, such that the frequency resolution can be improved for the purpose of noise reduction, while the original first-stage hearing-aid filter-bank system can still be used for compression and amplification. We also propose methods to implement the second filter-bank stage with little additional algorithmic delay. Furthermore, the computational complexity is an important design criterion. This finally leads to an application of the second filter-bank stage to lower frequency bands only, resulting in the ability to resolve the harmonics of speech. The paper presents a systematic description of the second filter-bank stage, discusses its influence on the processed signals in detail and further presents the results of a listening test which indicates the improved performance compared to the original single-stage filter-bank system. Alexander Schasse, Timo Gerkmann, Rainer Martin 0001, Wolfgang Sörgel, Thomas Pilgrim, Henning Puder |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Noise Power Spectral Density Estimation Using MaxNSR Blocking MatrixabstractIn this paper, a multi-microphone noise reduction system based on the generalized sidelobe canceller (GSC) structure is investigated. The system consists of a fixed beamformer providing an enhanced speech reference, a blocking matrix providing a noise reference by suppressing the target speech, and a single-channel spectral post-filter. The spectral post-filter requires the power spectral density (PSD) of the residual noise in the speech reference, which can in principle be estimated from the PSD of the noise reference. However, due to speech leakage in the noise reference, the noise PSD is overestimated, leading to target speech distortion. To minimize the influence of the speech leakage, a maximum noise-to-speech ratio (MaxNSR) blocking matrix is proposed, which maximizes the ratio between the noise and the speech leakage in the noise reference. The proposed blocking matrix can be computed from the generalized eigenvalue decomposition of the correlation matrix of the microphone signals and the noise coherence matrix, which is assumed to be time-invariant. Experimental results in both stationary and nonstationary diffuse noise fields show that the proposed algorithm outperforms existing blocking matrices in terms of target speech blocking ability, noise estimation and noise reduction performance. Timo Gerkmann, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | MMSE-optimal enhancement of complex speech coefficients with uncertain prior knowledge of the clean speech phaseabstractIn most STFT-based speech enhancement algorithms only the STFT amplitude of speech is processed, while the STFT phase of the noisy signal is neither modified nor employed to improve amplitude estimation. This is also, because modifying the spectral phase often yields undesired artifacts and unnatural sounding speech. In this paper, we first obtain a clean speech phase estimate using a recent phase reconstruction algorithm. Then, we propose to treat this reconstructed phase as uncertain a priori knowledge when deriving a joint MMSE estimate of the clean speech amplitude and phase. The resulting MMSE-estimator yields a compromise between the phase of the noisy signal and the prior phase estimate. Instrumental measures and informal listening show that the proposed estimator reduces un-desired artifacts and results in an improved speech quality. Timo Gerkmann |
ICASSP | 1 |
| 2014 | Frequency-domain single-channel inverse filtering for speech dereverberation: Theory and practiceabstractThe objective of single-channel inverse filtering is to design an inverse filter that achieves dereverberation while being robust to an inaccurate room impulse response (RIR) measurement or estimate. Since a stable and causal inverse filter typically does not exist, approximate time-domain inverse filtering techniques such as singlechannel least-squares (SCLS) have been proposed. However, besides being computationally expensive and often infeasible, SCLS generally leads to distortions in the output signal in the presence of RIR inaccuracies. In this paper, a theoretical analysis is initially provided, showing that the direct inversion of the acoustic transfer function in the frequency-domain generally yields instability and acausality issues. In order to resolve these issues, a novel frequency-domain inverse filtering technique is proposed that incorporates regularization and uses a single-channel speech enhancement scheme. Experimental results demonstrate that the proposed technique yields a higher dereverberation performance and has a significantly lower computational complexity compared to the SCLS technique. Ina Kodrasi, Timo Gerkmann, Simon Doclo |
ICASSP | 2 |
| 2014 | A posteriori voiced/unvoiced probability estimation based on a sinusoidal modelabstractIn this paper, we focus on methods for estimating the a posteriori probability of a signal segment being voiced which employ a harmonic signal model. Fisher et al. [1] present two likelihood functions for voiced and unvoiced speech from which the posterior probability can be derived. However, due to the chosen models, the a posteriori probability of a signal segment being voiced does not go to 0 % in unvoiced speech. Thus, a novel algorithm is proposed, which incorporates the expected unvoiced speech energy and allows for obtaining low probabilities. Further, it explicitly models the statistics of the segment energy and employs a state-of-the-art noise tracker. Experiments which were conducted on the TIMIT database for different noise types and noise levels show that the proposed method results in lower over-estimation and under-estimation of the voicing probability as compared to [1]. Robert Rehr, Martin Krawczyk-Becker, Timo Gerkmann |
ICASSP | 3 |
| 2014 | STFT phase reconstruction in voiced speech for an improved single-channel speech enhancementabstractThe enhancement of speech which is corrupted by noise is commonly performed in the short-time discrete Fourier transform domain. In case only a single microphone signal is available, typically only the spectral amplitude is modified. However, it has recently been shown that an improved spectral phase can as well be utilized for speech enhancement, e.g., for phase-sensitive amplitude estimation. In this paper, we therefore present a method to reconstruct the spectral phase of voiced speech from only the fundamental frequency and the noisy observation. The importance of the spectral phase is highlighted and we elaborate on the reason why noise reduction can be achieved by modifications of the spectral phase. We show that, when the noisy phase is enhanced using the proposed phase reconstruction, instrumental measures predict an increase of speech quality over a range of signal to noise ratios, even without explicit amplitude enhancement. Martin Krawczyk-Becker, Timo Gerkmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | On the relation between speech corruption models in the spectral and the cepstral domainabstractThe Gaussian distortion model in the short-time Fourier transform (STFT) domain is the basis of many of the modern speech enhancement algorithms. One of the reasons is that additive sources and late reverberation can be analyzed and processed quite efficiently in this domain. The STFT domain is however not well related to acoustic quality and is also not well suited for learning models due to the high variability of speech in this domain. On the other hand, the cepstral domain has proved to be very well suited for these last two purposes, however, at the cost of loosing the simple linear relation between desired source and additive interferences. In this paper we explore the relation between the Gaussian distortion models in the STFT and the cepstral domain. We show how the assumption of a jointly Gaussian distortion model in the cepstrum domain is fulfilled for well-known distortion models in STFT domain. We provide closed-form solutions relating the joint distributions of corrupted and clean speech in the STFT and the cepstrum domain. We also propose various ways in which this model can be used to enhance speech. Ramón Fernandez Astudillo, Timo Gerkmann |
ICASSP | 2 |
| 2013 | Privacy-preserving distributed speech enhancement forwireless sensor networks by processing in the encrypted domainabstractTo improve speech communication in noisy and reverberant environments, an increased interest is shown to develop algorithms that make efficiently use of acoustic wireless sensor networks (WSNs). The processors and sensors forming these WSNs can be owned by multiple users. Sending private data across such a WSN can lead to severe privacy and security issues and may limit its acceptance. Using the advantages of WSNs, while guaranteeing people's privacy, requires therefore to share processors and data in a privacy preserving manner. In this paper we raise attention to the problem of privacy and security for distributed speech enhancement and propose the new paradigm of privacy preserving distributed beamforming. Using cryptographic techniques, particularly homomorphic encryption, we demonstrate how distributed beamforming techniques can be computed in a privacy preserving manner in the encrypted domain. Richard C. Hendriks, Zekeriya Erkin, Timo Gerkmann |
ICASSP | 3 |
| 2013 | MMSE-Optimal Spectral Amplitude Estimation Given the STFT-PhaseabstractIn this letter, we derive a minimum mean squared error (MMSE) optimal estimator for clean speech spectral amplitudes, which we apply in single channel speech enhancement. As opposed to state-of-the-art estimators, the optimal estimator is derived for a given clean speech spectral phase. We show that the phase contains additional information that can be exploited to distinguish outliers in the noise from the target signal. With the proposed technique, incorporating the phase can potentially improve the PESQ-MOS by 0.5 in babble noise as compared to state-of-the-art amplitude estimators. In a blind setup we achieve a PESQ improvement of around 0.25 in voiced speech. Timo Gerkmann, Martin Krawczyk-Becker |
IEEE Signal Process. Lett. | 1 |
| 2012 | Improved mmse-based noise PSD tracking using temporal cepstrum smoothingabstractRecently, it has been shown that MMSE-based noise power estimation [1] results in an improved noise tracking performance with respect to minimum statistics-based approaches. The MMSE-based approach employs two estimates of the speech power to estimate the unbiased noise power. In this work, we improve the MMSE-based noise power estimator by employing a more advanced estimator of the speech power based on temporal cepstrum smoothing (TCS). TCS can exploit knowledge about the speech spectral structure. As a result, only one speech power estimate is needed for MMSE-based noise power estimation. Moreover, the presented estimator results in an improved noise tracking performance, especially in babble noise, where SNR improvements of 1dB over the original MMSE-based approach can be observed. Timo Gerkmann, Richard C. Hendriks |
ICASSP | 1 |
| 2012 | Unbiased MMSE-Based Noise Power Estimation With Low Complexity and Low Tracking DelayabstractRecently, it has been proposed to estimate the noise power spectral density by means of minimum mean-square error (MMSE) optimal estimation. We show that the resulting estimator can be interpreted as a voice activity detector (VAD)-based noise power estimator, where the noise power is updated only when speech absence is signaled, compensated with a required bias compensation. We show that the bias compensation is unnecessary when we replace the VAD by a soft speech presence probability (SPP) with fixed priors. Choosing fixed priors also has the benefit of decoupling the noise power estimator from subsequent steps in a speech enhancement framework, such as the estimation of the speech power and the estimation of the clean speech. We show that the proposed speech presence probability (SPP) approach maintains the quick noise tracking performance of the bias compensated minimum mean-square error (MMSE)-based approach while exhibiting less overestimation of the spectral noise power and an even lower computational complexity. Timo Gerkmann, Richard C. Hendriks |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Noise Correlation Matrix Estimation for Multi-Microphone Speech EnhancementabstractFor multi-channel noise reduction algorithms like the minimum variance distortionless response (MVDR) beamformer, or the multi-channel Wiener filter, an estimate of the noise correlation matrix is needed. For its estimation, it is often proposed in the literature to use a voice activity detector (VAD). However, using a VAD the estimated matrix can only be updated in speech absence. As a result, during speech presence the noise correlation matrix estimate does not follow changing noise fields with an appropriate accuracy. This effect is further increased, as in nonstationary noise voice activity detection is a rather difficult task, and false-alarms are likely to occur. In this paper, we present and analyze an algorithm that estimates the noise correlation matrix without using a VAD. This algorithm is based on measuring the correlation of the noisy input and a noise reference which can be obtained, e.g., by steering a null towards the target source. When applied in combination with an MVDR beamformer, it is shown that the proposed noise correlation matrix estimate results in a more accurate beamformer response, a larger signal-to-noise ratio improvement and a larger instrumentally predicted speech intelligibility when compared to competing algorithms such as the generalized sidelobe canceler, a VAD-based MVDR beamformer, and an MVDR based on the noisy correlation matrix. Richard C. Hendriks, Timo Gerkmann |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Estimation of the noise correlation matrixabstractTo harvest the potential of multi-channel noise reduction methods, it is crucial to have an accurate estimate of the noise correlation matrix. Existing algorithms either assume speech absence and exploit a voice activity detector (VAD), or make use of additional assumptions like a diffuse noise field. Therefore, these algorithms are limited with respect to their tracking speed and the type of noise fields for which they can estimate the correlation matrix. In this paper we present a new method for noise correlation matrix estimation that makes no assumptions about the type of noise field, nor uses a VAD. The presented method exploits the existence of accurate single-channel noise PSD estimators, as well as the avail ability of one noise reference per microphone pair. For spatially and temporally non-stationary noise fields, the proposed method leads to improved performance compared to widely used state-of-the-art reference methods in terms of both segmental SNR and beamformer response error. Richard C. Hendriks, Timo Gerkmann |
ICASSP | 2 |
| 2010 | Speech presence probability estimation based on temporal cepstrum smoothingabstractWe propose a novel, robust estimator for the probability of speech presence at each time-frequency point in the short-time discrete Fourier domain. While existing estimators perform quite reliably in stationary noise environments, they usually exhibit a large false-alarm rate in nonstationary noise that results in a great deal of noise leakage when applied to a speech enhancement task. The proposed estimator overcomes this problem by temporally smoothing the cepstrum of the a posteriori signal-to-noise ratio (SNR), and yields considerably less noise leakage and low speech distortions in both, stationary and nonstationary noise as compared to state-of-the-art estimators. Especially in babble noise, this results in large SNR improvements. Timo Gerkmann, Martin Krawczyk-Becker, Rainer Martin 0001 |
ICASSP | 1 |
| 2009 | Multi-microphone maximum a posteriori fundamental frequency estimation in the cepstral domainabstractIn this work we derive a new cepstrum based maximum likelihood fundamental frequency estimator that exploits the information of multiple microphones. The new approach results in a maximum search on the sum of the microphone cepstra. We compare the new approach to a maximum search on the cepstrum of the output signal of a delay-and-sum beamformer. We show that the new approach outperforms the beamforming approach for all considered input signal-to-noise ratios. We develop a general framework which includes the cepstral harmonics of the fundamental frequency and extend the approach towards a maximum a posteriori fundamental period tracker that further enhances the results and increases the robustness in noisy environments. Timo Gerkmann, Rainer Martin 0001, Derya Dalga |
ICASSP | 1 |
| 2008 | A novel a priori SNR estimation approach based on selective cepstro-temporal smoothingabstractWhile state-of-the-art approaches obtain an estimate of the a priori SNR by adaptively smoothing its maximum likelihood estimate in the frequency domain, we selectively smooth the maximum likelihood estimate in the cepstral domain. In the cepstral domain the noisy speech signal is decomposed into coefficients related mainly to the speech envelope, the excitation, and noise. As in the cepstral domain coefficients that represent speech can be robustly determined, we can apply little smoothing to speech coefficients and strong smoothing to noise coefficients. Thus, speech components are preserved and musical noise is suppressed. In speech enhancement experiments we obtain consistent improvements over the well known decision-directed approach. Colin Breithaupt, Timo Gerkmann, Rainer Martin 0001 |
ICASSP | 2 |
| 2008 | Improved A Posteriori Speech Presence Probability Estimation Based on a Likelihood Ratio With Fixed PriorsabstractIn this paper, we present an improved estimator for the speech presence probability at each time-frequency point in the short-time Fourier transform domain. In contrast to existing approaches, this estimator does not rely on an adaptively estimated and thus signal-dependent a priori signal-to-noise ratio estimate. It therefore decouples the estimation of the speech presence probability from the estimation of the clean speech spectral coefficients in a speech enhancement task. Using both a fixed a priori signal-to-noise ratio and a fixed prior probability of speech presence, the proposed a posteriori speech presence probability estimator achieves probabilities close to zero for speech absence and probabilities close to one for speech presence. While state-of-the-art speech presence probability estimators use adaptive prior probabilities and signal-to-noise ratio estimates, we argue that these quantities should reflect true a priori information that shall not depend on the observed signal. We present a detection theoretic framework for determining the fixed a priori signal-to-noise ratio. The proposed estimator is conceptually simple and yields a better tradeoff between speech distortion and noise leakage than state-of-the-art estimators. Timo Gerkmann, Colin Breithaupt, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Cepstral Smoothing of Spectral Filter Gains for Speech Enhancement Without Musical NoiseabstractMany speech enhancement algorithms that modify short-term spectral magnitudes of the noisy signal by means of adaptive spectral gain functions are plagued by annoying spectral outliers. In this letter, we propose cepstral smoothing as a solution to this problem. We show that cepstral smoothing can effectively prevent spectral peaks of short duration that may be perceived as musical noise. At the same time, cepstral smoothing preserves speech onsets, plosives, and quasi-stationary narrowband structures like voiced speech. The proposed recursive temporal smoothing is applied to higher cepstral coefficients only, excluding those representing the pitch information. As the higher cepstral coefficients describe the finer spectral structure of the Fourier spectrum, smoothing them along time prevents single coefficients of the filter function from changing excessively and independently of their neighboring bins, thus suppressing musical noise. The proposed cepstral smoothing technique is very effective in nonstationary noise. Colin Breithaupt, Timo Gerkmann, Rainer Martin 0001 |
IEEE Signal Process. Lett. | 2 |
| 2006 | Statistical Inference of Missing Speech Data in the ICA DomainabstractWe address the problem of speech estimation as statistical estimation with "missing" data in the independent component analysis (ICA) domain. Missing components are substituted by values drawn from "similar" data in a multi-faceted ICA representation of the complete data. The paper presents the algorithm for the inference of missing data in the case of a fixed pattern of missing data. We apply our approach to the problem of bandwidth extension, or where speech is degraded by a fixed filtering process and show the capability of the algorithm to reconstruct fine missing details of the original data with little artifacts. The evaluation is done using objective distortion measures on speech samples from the NTT database Justinian P. Rosca, Timo Gerkmann, Doru-Cristian Balcan |
ICASSP (5) | 2 |
| 2006 | Soft decision combining for dual channel noise reduction
Timo Gerkmann, Rainer Martin 0001 |
INTERSPEECH | 1 |