EDBT 2026 Demo / reviewers in the wild / expert
Sebastian Braun
dblp:136/5248
· DBLP profile ↗
29ranked-venue papers
10as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ultra-Low Latency Speech Enhancement -A Comprehensive StudyabstractSpeech enhancement models should meet very low latency requirements typically smaller than 5 ms for hearing assistive devices. While various low-latency techniques have been proposed, comparing these methods in a controlled setup using DNNs remains blank. Previous papers have variations in task, training data, scripts, and evaluation settings, which make fair comparison impossible. Moreover, all methods are tested on small, simulated datasets, making it difficult to fairly assess their performance in real-world conditions, which could impact the reliability of scientific findings. To address these issues, we comprehensively investigate various low-latency techniques using consistent training on large-scale data and evaluate with more relevant metrics on real-world data. Specifically, we explore the effectiveness of asymmetric windows, learnable windows, adaptive time domain filterbanks, and the future-frame prediction technique. Additionally, we examine whether increasing the model size can compensate for the reduced window size. Sebastian Braun |
ICASSP | 2 |
| 2024 | Adapting Frechet Audio Distance for Generative Music EvaluationabstractThe growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual quality is understudied. We show that FAD performance may be hampered by sample size bias, poor choice of audio embeddings, or the use of biased or low-quality reference sets. We propose reducing sample size bias by extrapolating scores towards an infinite sample size. Through comparisons with MusicCaps labels and a listening test we identify audio embeddings and music reference sets that yield FAD scores well-correlated with acoustic and musical quality. Our results suggest that per-song FAD can be useful to identify outlier samples and predict perceptual quality for a range of music sets and generative models. Finally, we release a toolkit that allows adapting FAD for generative music evaluation. Azalea Gui, Hannes Gamper, Sebastian Braun, Dimitra Emmanouilidou |
ICASSP | 3 |
| 2024 | Visualizing Uncertainty in AI for Accident Severity ClassificationabstractUncertainties are frequently encountered in machine learning, especially in classification tasks where the model's output is determined by the probabilities of each class. In geospatial use cases, the challenge is to find appropriate visualizations on geographic maps for these uncertainties, so that domain experts know when to trust the model. This study utilizes the uncertainties of an extreme gradient boosting machine to classify traffic accident severity and presents various visualization profiles. These visualizations are the outcome of an experimental approach. Each profile has its own advantages but may also encounter issues such as overlapping data points or the determination of what uncertain means. The most frequently used visualization types are symbols and colors. Cédric Roussel, Klaus Böhm, Bastian Jakobi, Alisa Vlasov, Sebastian Braun, Alexander Rolwes |
IV | 6 |
| 2023 | Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant EnvironmentsabstractReal-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberated and noise-free speech tracks, each track containing only one talker. While large state-of-the-art DNNs can achieve excellent separation from anechoic mixtures of speech, the main challenge is to create compact and causal models that can separate reverberant mixtures at inference time. In this paper, we explore low-complexity, resource-efficient, causal DNN architectures for real-time separation of two or more simultaneous speakers. A cascade of three neural network modules are trained to sequentially perform noise-suppression, separation, and de-reverberation. For comparison, a larger end-to-end model is trained to output two anechoic speech signals directly from noisy reverberant speech mixtures. We propose an efficient single-decoder architecture with "subtractive" separation for real-time recursive speech separation for two or more speakers. Evaluation on real monophonic recordings of speech mixtures, according to speech separation measures like SI-SDR, perceptual measures like DNS-MOS, and a novel proposed channel separation metric, show that these compact causal models can separate speech mixtures with low latency, and perform on par with large offline state-of-the-art models like SepFormer. Julian Neri, Sebastian Braun |
ICASSP | 2 |
| 2022 | Effect of Noise Suppression Losses on Speech Distortion and ASR PerformanceabstractDeep learning based speech enhancement has made rapid development towards improving quality, while models are becoming more compact and usable for real-time on-the-edge inference. However, the speech quality scales directly with the model size, and small models are often still unable to achieve sufficient quality. Furthermore, the introduced speech distortion and artifacts greatly harm speech quality and intelligibility, and often significantly degrade automatic speech recognition (ASR) rates. In this work, we shed light on the success of the spectral complex compressed mean squared error (MSE) loss, and how its magnitude and phase-aware terms are related to the speech distortion vs. noise reduction trade off. We further investigate integrating pre-trained reference-less predictors for mean opinion score (MOS) and word error rate (WER), and pre-trained embeddings on ASR and sound event detection. Our analyses reveal that none of the pre-trained networks added significant performance over the strong spectral loss. Sebastian Braun, Hannes Gamper |
ICASSP | 1 |
| 2022 | ICASSP 2022 Acoustic Echo Cancellation ChallengeabstractThe ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios, adding speech recognition word accuracy rate as a metric, and making the audio 48 kHz. We open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 10,000 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source an online subjective test framework and provide an online objective metric service for researchers to quickly test their results. The winners of this challenge were selected based on the average Mean Opinion Score (MOS) achieved across all scenarios and the word accuracy rate. Ross Cutler, Ando Saabas, Tanel Pärnamaa, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner |
ICASSP | 6 |
| 2022 | Icassp 2022 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-source datasets and test sets for researchers to train their deep noise suppression models, as well as a subjective evaluation framework based on ITU-T P.835 to rate and rank-order the challenge entries. We provide access to DNS-MOS P.835 and word accuracy (WAcc) APIs to challenge participants to help with iterative model improvements. In this challenge, we introduced the following changes: (i) Included mobile device scenarios in the blind test set; (ii) Included a personalized noise suppression track with baseline; (iii) Added WAcc as an objective metric; (iv) Included DNSMOS P.835; (v) Made the training datasets and test sets fullband (48 kHz). We use an average of WAcc and subjective scores P.835 SIG, BAK, and OVRL to get the final score for ranking the DNS models. We believe that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-world scenarios. Harishchandra Dubey, Vishak Gopal, Ross Cutler, Ashkan Aazami, Sergiy Matusevych, Sebastian Braun, Sefik Emre Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, Robert Aichner |
ICASSP | 6 |
| 2022 | Unsupervised Speech Enhancement with Speech Recognition Embedding and Disentanglement LossesabstractSpeech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority of training datasets for speech enhancement systems are synthetic. When mixing clean speech and noisy corpora to create the synthetic datasets, domain mismatches occur between synthetic and real-world recordings of noisy speech or audio. Second, there is a trade-off between increasing speech enhancement performance and degrading speech recognition (ASR) performance. Thus, we propose an unsupervised loss function to tackle those two problems. Our function is developed by extending the MixIT loss function with speech recognition embedding and disentanglement loss. Our results show that the proposed function effectively improves the speech enhancement performance compared to a baseline trained in a supervised way on the noisy VoxCeleb dataset. While fully unsupervised training is unable to exceed the corresponding baseline, with joint super- and unsupervised training, the system is able to achieve similar speech quality and better ASR performance than the best supervised baseline. Viet Anh Trinh, Sebastian Braun |
ICASSP | 2 |
| 2022 | Performance optimizations on U-Net speech enhancement modelsabstractDeep learning approaches-while remarkably successful in audio enhancement-result in slower inference times which are prohibitive in real-time applications. We develop a practical strategy for compressing U-Net style deep neural network architectures. On deep noise suppression (DNS) models we achieve a state-of-the-art 7.25× inference speed up over the baseline CRUSE model, with a smooth model performance degradation. Our method is friendly to practitioners as it only requires setting a single compression parameter, while achieving non-uniform compression rates across layers. We report inference speed because a parameter or memory reduction does not necessitate speedup, and we measure model quality using an accurate non-intrusive objective speech quality metric. Jerry Chee, Sebastian Braun, Vishak Gopal, Ross Cutler |
MMSP | 2 |
| 2022 | Gamification of virtual reality assembly training: Effects of a combined point and level system on motivation and training results
Jessica Ulmer, Sebastian Braun, Chi-Tsun Cheng, Steve Dowey, Jörg F. Wollert |
Int. J. Hum. Comput. Stud. | 2 |
| 2021 | DBnet: Doa-Driven Beamforming Network for end-to-end Reverberant Sound Source SeparationabstractMany deep learning techniques are available to perform source separation and reduce background noise. However, designing an end-to-end multi-channel source separation method using deep learning and conventional acoustic signal processing techniques still remains challenging. In this paper we propose a direction-of-arrival-driven beamforming network (DBnet) consisting of direction-of-arrival (DOA) estimation and beamforming layers for end-to-end source separation. We propose to train DBnet using loss functions that are solely based on the distances between the separated speech signals and the target speech signals, without a need for the ground-truth DOAs of speakers. To improve the source separation performance, we also propose end-to-end extensions of DBnet which incorporate post masking networks. We evaluate the proposed DBnet and its extensions on a very challenging dataset, targeting realistic far-field sound source separation in reverberant and noisy environments. The experimental results show that the proposed extended DBnet using a convolutional-recurrent post masking network outperforms state-of-the-art source separation methods. Ali Aroudi, Sebastian Braun |
ICASSP | 2 |
| 2021 | Towards Efficient Models for Real-Time Deep Noise SuppressionabstractWith recent research advancements, deep learning models are be-coming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality and background noise reduction, the main challenge is to obtain compact enough models, which are resource efficient during inference time. An important but often neglected aspect for data-driven methods is that results can be only convincing when tested on real-world data and evaluated with useful metrics. In this work, we investigate reasonably small recurrent and convolutional-recurrent network architectures for speech enhancement, trained on a large dataset considering also reverberation. We show interesting tradeoffs between computational complexity and the achievable speech quality, measured on real recordings using a highly accurate MOS estimator. It is shown that the achievable speech quality is a function of network complexity, and show which models have better tradeoffs. Sebastian Braun, Hannes Gamper, Chandan K. A. Reddy, Ivan Tashev |
ICASSP | 1 |
| 2021 | ICASSP 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to train their noise suppression models. We also open-sourced a subjective evaluation framework and used the tool to evaluate and select the final winners. Many researchers from academia and industry made significant contributions to push the field forward. We also learned that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-time conditions. In this challenge, we expanded both our training and test datasets. Clean speech in the training set has increased by 200% with the addition of singing voice, emotion data, and non-English languages. The test set has increased by 100% with the addition of singing, emotional, non-English (tonal and non-tonal) languages, and, personalized DNS test clips. There are two tracks with focus on (i) real-time denoising, and (ii) real-time personalized DNS. We present the challenge results at the end. Chandan K. A. Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
ICASSP | 5 |
| 2021 | ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and ResultsabstractThe ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 2,500 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source two large test sets, and we open source an online subjective test framework for researchers to quickly test their results. The winners of this challenge will be selected based on the average Mean Opinion Score (MOS) achieved across all different single talk and double talk scenarios. Kusha Sridhar, Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Hannes Gamper, Sebastian Braun, Robert Aichner, Sriram Srinivasan 0003 |
ICASSP | 7 |
| 2021 | INTERSPEECH 2021 Acoustic Echo Cancellation Challenge
Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Sten Sootla, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 8 |
| 2021 | INTERSPEECH 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP 2020. We open-sourced training and test datasets for the wideband scenario. We also open-sourced a subjective evaluation framework based on ITU-T standard P.808, which was also used to evaluate participants of the challenge. Many researchers from academia and industry made significant contributions to push the field forward, yet even the best noise suppressor was far from achieving superior speech quality in challenging scenarios. In this version of the challenge organized at INTERSPEECH 2021, we are expanding both our training and test datasets to accommodate full band scenarios. The two tracks in this challenge will focus on real-time denoising for (i) wide band, and(ii) full band scenarios. We are also making available a reliable non-intrusive objective speech quality metric called DNSMOS for the participants to use during their development phase. Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 7 |
| 2020 | Predicting Word Error Rate for Reverberant SpeechabstractReverberation negatively impacts the performance of automatic speech recognition (ASR). Prior work on quantifying the effect of reverberation has shown that clarity (C50), a parameter that can be estimated from the acoustic impulse response, is correlated with ASR performance. In this paper we propose predicting ASR performance in terms of the word error rate (WER) directly from acoustic parameters via a polynomial, sigmoidal, or neural network fit, as well as blindly from reverberant speech samples using a convolutional neural network (CNN). We carry out experiments on two state-of-the-art ASR models and a large set of acoustic impulse responses (AIRs). The results confirm C50 and C80 to be highly correlated with WER, allowing WER to be predicted with the proposed fitting approaches. The proposed non-intrusive CNN model outperforms C50-based WER prediction, indicating that WER can be estimated blindly, i.e., directly from the reverberant speech samples without knowledge of the acoustic parameters. Hannes Gamper, Dimitra Emmanouilidou, Sebastian Braun, Ivan Tashev |
ICASSP | 3 |
| 2020 | Joint Beamforming and Reverberation Cancellation Using a Constrained Kalman Filter With Multichannel Linear PredictionabstractThe performance of speech processing systems degrades significantly in far-field scenarios where the distance between the user and microphones increases, leading to low signal-to-noise and signal-to-reverberation ratios. To address this challenge, combining the denoising and dereverberation techniques in both parallel and cascade configurations has been widely studied. However, a parallel or cascade combination may not be efficient while imposing a large computational complexity. We propose a constrained Kalman filter based multichannel linear prediction method to jointly perform denoising and dereverberation efficiently using an online processing algorithm. In contrast to previously proposed methods which utilize steering vectors based on the relative early transfer function, our algorithm is implemented using a direct relative transfer function based steering vector, which aims at extracting the direct sound as opposed to preserving the early reflections. We show that the proposed algorithm outperforms existing online implementations of integrated beamformer and linear prediction methods on the REVERB challenge speech enhancement task while being computationally less complex. Sahar Hashemgeloogerdi, Sebastian Braun |
ICASSP | 2 |
| 2020 | Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech EnhancementabstractThis paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that enhances short-time speech spectra on a single-frame-in, single-frame-out basis, a framework adopted by most classical signal processing methods. We propose two novel mean-squared-error-based learning objectives that enable separate control over the importance of speech distortion versus noise reduction. The proposed loss functions are evaluated by widely accepted objective quality and intelligibility measures and compared to other competitive online methods. In addition, we study the impact of feature normalization and varying batch sequence lengths on the objective quality of enhanced speech. Finally, we show subjective ratings for the proposed approach and a state-of-the-art real-time RNN-based method. Yangyang Xia, Sebastian Braun, Chandan K. A. Reddy, Harishchandra Dubey, Ross Cutler, Ivan Tashev |
ICASSP | 2 |
| 2020 | The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge ResultsabstractThe INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical approach to evaluate the noise suppression methods is to use objective metrics on the test set obtained by splitting the original dataset. While the performance is good on the synthetic test set, often the model performance degrades significantly on real recordings. Also, most of the conventional objective metrics do not correlate well with subjective tests and lab subjective tests are not scalable for a large test set. In this challenge, we open-sourced a large clean speech and noise corpus for training the noise suppression models and a representative test set to real-world scenarios consisting of both synthetic and real recordings. We also open-sourced an online subjective test framework based on ITU-T P.808 for researchers to reliably test their developments. We evaluated the results using P.808 on a blind test set. The results and the key learnings from the challenge are discussed. The datasets and scripts can be found here for quick access https://github.com/microsoft/DNS-Challenge. Chandan K. A. Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, Puneet Rana, Sriram Srinivasan 0003, Johannes Gehrke |
INTERSPEECH | 10 |
| 2019 | Directional Interference Suppression Using a Spatial Relative Transfer Function FeatureabstractMany speech enhancement systems consist of a beamformer and a spectral suppression postfilter. While it is well understood how to design beamformers to suppress either non-directional or directional interference, their suppression ability is limited by the number, type and positions of microphones. However, spectral postfilters that can further increase the suppression, are usually only designed to suppress non-directional noise. In this work, we propose a spatially selective spectral suppressor addressing directional and non-directional interference. The proposed suppressor is based on the relative transfer function of the target source location. While existing directional suppression techniques are limited to farfield scenarios or certain microphone geometries, we propose a general approach without restrictions on the microphone array and without farfield assumption. We show that the proposed spatial suppressor is able to suppress noise and directional interfering speakers, which substantially improves the performance of speech recognizer, and reduces undesired recognition of interfering talkers. Sebastian Braun, Ivan Tashev |
ICASSP | 1 |
| 2018 | Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio EstimationabstractNon-intrusive estimators for acoustic parameters like the direct-to-reverberation ratio (DRR) are useful tools but still perform weakly as shown in the acoustic characterization of environments (ACE) challenge. In this paper, we develop a novel dual-channel metric based on the modulation energy domain for DRR estimation. In contrast to established modulation based single-channel metrics like the speech-to-reverberation modulation energy ratio (SRMR), we exploit the spatial information from two microphones as well as the temporal dynamics in the modulation energy domain. The developed metric shows a strong linear correlation to the DRR, which allows a simple mapping. It is shown that the metric is robust against the microphone array configuration, room characteristics and the speech signal. The proposed metric is compared to a reference method based on the spectral variance of the room transfer functions, and both metrics are evaluated using simulated and measured data. In our experiments, the proposed metric achieved a higher correlation and lower RMSE compared the reference method, and outperforms existing SRMR based DRR estimators. Sebastian Braun, João Felipe Santos, Emanuël A. P. Habets, Tiago H. Falk |
ICASSP | 1 |
| 2018 | Single-Channel Dereverberation Using Direct MMSE Optimization and Bidirectional LSTM NetworksabstractDereverberation is useful in hands-free communication and voice controlled devices for distant speech acquisition. Single-channel dereverberation can be achieved by applying a time-frequency (TF) mask to the short-time Fourier transform (STFT) representation of a reverberant signal. Recent approaches have used deep neural networks (DNNs) to estimate such masks. Previously proposed DNN-based mask estimation methods train a DNN to minimize the mean-squared-error (MSE) between the desired and estimated masks. Recent TF mask estimation methods for signal separation directly minimize instead the MSE between the desired and estimated STFT magnitudes. We apply this direct optimization concept to dereverberation. Moreover, as reverberation exceeds the duration of a single STFT frame, we propose to use a bidirectional long short-term memory (LSTM) network which is able to take the relation between multiple STFT frames into account. We evaluated our method for different reverberation times and source-microphone distances using simulated as well as measured room impulse responses of different rooms. An evaluation of the proposed method and a comparison with a state-of-the-art method demonstrate the superiority of our approach and its robustness to different acoustic conditions. Wolfgang Mack, Soumitro Chakrabarty, Fabian-Robert Stöter, Sebastian Braun, Bernd Edler, Emanuël A. P. Habets |
INTERSPEECH | 4 |
| 2018 | Linear Prediction-Based Online Dereverberation and Noise Reduction Using Alternating Kalman FiltersabstractMultichannel linear prediction-based dereverberation in the short-time Fourier transform (STFT) domain has been shown to be highly effective. Using this framework, the desired dereverberated multichannel signal is obtained by filtering the noise-free reverberant signals using the estimated multichannel autoregressive (MAR) coefficients. To use such methods in the presence of noise, especially in the case of online processing, remains a challenging problem. Existing sequential enhancement structures, which first remove the noise and then estimate the MAR coefficients, suffer from a causality problem as both the optimal noise reduction and dereverberation stages depend on the current output of each other. To address this problem, an algorithm that consists of two alternating Kalman filters to estimate the noise-free reverberant signals and the (MAR) coefficients is proposed. The causality of the estimation procedure is important when dealing with time-variant acoustic scenarios, where the MAR coefficients are time-varying. The proposed method is evaluated using simulated and measured acoustic impulse responses and is compared to a method based on the same signal model. In addition, a method to control the reverberation reduction and noise reduction independently is derived. Sebastian Braun, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Evaluation and Comparison of Late Reverberation Power Spectral Density EstimatorsabstractReduction of late reverberation can be achieved using spatio-spectral filters, such as the multichannel Wiener filter. To compute this filter, an estimate of the late reverberation power spectral density (PSD) is required. In recent years, a multitude of late reverberation PSD estimators have been proposed. In this paper, these estimators are categorized into several classes, their relations and differences are discussed, and a comprehensive experimental comparison is provided. To compare their performance, simulations in controlled as well as practical scenarios are conducted. It is shown that a common weakness of spatial coherence-based estimators is their performance in high direct-to-diffuse ratio conditions. To mitigate this problem, a correction method is proposed and evaluated. It is shown that the proposed correction method can decrease the speech distortion without significantly affecting the reverberation reduction. Sebastian Braun, Adam Kuklasinski, Ofer Schwartz, Oliver Thiergart, Emanuël A. P. Habets, Sharon Gannot, Simon Doclo, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Online Dereverberation for Dynamic Scenarios Using a Kalman Filter With an Autoregressive ModelabstractReverberant signals can be modeled in the short-time Fourier transform domain using a frequency-dependent autoregressive (AR) model. In state-of-the-art, these AR coefficients have been considered stationary, which does not hold in time-varying environments. We propose to model these AR coefficients using a first-order Markov process, whereas the reverberant microphone signal observations are considered deterministic. This leads to a framework where the AR coefficients can be optimally estimated using a Kalman filter per subband. As a consequence, we can dereverberate the observed signals by applying the estimated AR coefficients as an adaptive linear filter per subband. Estimators for the required statistical parameters in the Kalman filter are derived. Due to the adaptive solution, the algorithm is suitable for real-time applications. It is shown that the proposed method outperforms an existing recursive least-squares solution in terms of reverberation reduction, convergence time, and tracking changes in the acoustic scene. Sebastian Braun, Emanuël A. P. Habets |
IEEE Signal Process. Lett. | 1 |
| 2015 | Residual noise control using a parametric multichannel Wiener filterabstractMultichannel noise reduction techniques are commonly used in speech communication applications. In these applications, it is often desired to maintain a residual amount of background noise to avoid perceptually unpleasant artifacts, such as musical tones or time periods of complete silence. Noise reduction can be achieved by the parametric multichannel Wiener filter (PMWF), which provides a trade-off between speech distortion and noise reduction. To additionally control the maximum noise reduction, the PMWF can be decomposed into a spatial filter and a spectral gain, which is limited to a desired minimum value. Such decomposition is however only possible if the desired source power spectral density matrix is rank-one, which in general does not even hold for a single source in reverberant environments. In the proposed approach, we define the desired signal as a sum of the speech signal plus the desired residual noise, and derive an optimum filter in the minimum mean-square error sense. The resulting filter has the advantage that it enables direct control of the maximum noise reduction without the need for a gain limiting step and is furthermore applicable to desired signals of higher rank. We analyze the derived filter thoroughly and show its relation to the standard PMWF that results as a special case. Furthermore, we propose a solution for keeping the residual noise level constant in slowly time-varying noise fields. Sebastian Braun, Konrad Kowalczyk, Emanuël A. P. Habets |
ICASSP | 1 |
| 2014 | Automatic spatial gain control for an informed spatial filterabstractWhen capturing speech in a multi-talker telecommunication scenario, it is desirable to keep the enhanced signal at an equal loudness level for each speaker. Single-channel automatic gain control systems are not able to adjust the level of different talkers when they are simultaneously active. In this work, an automatic spatial gain control (ASGC) algorithm is proposed that adjusts the directional response of an existing informed spatial filter such that the direct sound of multiple sources can be kept at a constant desired loudness level at the output. The spatial filter additionally reduces diffuse sound and ambient noise. It is shown that the proposed AGSC works well within the tested scenario, and is able to adjust the levels of different speakers even during double talk scenarios. Sebastian Braun, Oliver Thiergart, Emanuël A. P. Habets |
ICASSP | 1 |
| 2013 | An informed spatial filter for dereverberation in the spherical harmonic domainabstractIn speech communication systems the received microphone signals are commonly degraded by reverberation and ambient noise that can decrease the fidelity and intelligibility of a desired speaker. Reverberation can be modeled as non-stationary diffuse sound which is not directly observable. In this work, we derive a multichannel Wiener filter in the spherical harmonic domain to reduce both reverberation and noise. The filter depends on the direction-of-arrival of the direct sound of the desired speaker and an interference power spectral density matrix for which an estimator is developed. The resulting informed spatial filter incorporates instantaneous information about the diffuseness of the sound field into the design of the filter. In addition, it is shown how the proposed filter relates to the well-known robust minimum variance distortionless response filter that is also used for comparison in the evaluation. Experimental results show that the proposed spatial filter provides a tradeoff between noise reduction and dereverberation depending on the diffuse sound PSD. Sebastian Braun, Daniel P. Jarrett, Johannes Fischer 0002, Emanuël A. P. Habets |
ICASSP | 1 |