VLDB 2026 Research / reviewers in the wild / expert
Simon Doclo
dblp:11/1434
· DBLP profile ↗
116ranked-venue papers
10as first author
21since 2021 · last 2025
0000-0002-3392-2381ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 47 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple SpeakersabstractTo estimate the direction of arrival (DOA) of multiple speakers, subspace-based prototype transfer function matching methods such as multiple signal classification (MUSIC) or relative transfer function (RTF) vector matching are commonly employed. In general, these methods require calibrated microphone arrays, which are characterized by a known array geometry or a set of known prototype transfer functions for several directions. In this paper, we consider a partially calibrated microphone array, composed of a calibrated binaural hearing aid and a (non-calibrated) external microphone at an unknown location with no available set of prototype transfer functions. We propose a procedure for completing sets of prototype transfer functions by exploiting the orthogonality of subspaces, allowing to apply matching-based DOA estimation methods with partially calibrated microphone arrays. For the MUSIC and RTF vector matching methods, experimental results for two speakers in noisy and reverberant environments clearly demonstrate that for all locations of the external microphone DOAs can be estimated more accurately with completed sets of prototype transfer functions than with incomplete sets. Daniel Fejgin, Simon Doclo |
ICASSP | 2 |
| 2025 | Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear MicrophoneabstractHearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user’s own voice in a noisy environment. In this scenario, own voice reconstruction (OVR) is essential for enhancing the quality and intelligibility of the recorded noisy own voice signals for telephony applications. In previous work, we developed a deep learning-based OVR system, aiming to reduce the amount of device-specific recorded signals for training by using data augmentation with phoneme-dependent models of own voice transfer characteristics. Given the limited computational resources available on hearables, in this paper we propose low-complexity variants of an OVR system based on the frequency and time joint non-linear filter (FT-JNF) architecture and investigate the required amount of device-specific recorded signals for effective data augmentation and fine-tuning. Simulation results show that the proposed OVR system considerably improves speech quality, even under constraints of low complexity and a limited amount of device-specific recorded signals. Mattes Ohlenbusch, Christian Rollwage, Simon Doclo |
ICASSP | 3 |
| 2024 | Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple SpeakersabstractTo estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different frequencies can be combined. According to how the SPS are combined, frequency fusion mechanisms are categorized into narrowband, broadband, or speaker-grouped, where the latter mechanism requires a speaker-wise grouping of frequencies. For a binaural hearing aid setup, in this paper we propose an interaural time difference (ITD)-based speaker-grouped frequency fusion mechanism. By exploiting the DOA dependence of ITDs, frequencies can be grouped according to a common ITD and be used for DOA estimation of the respective speaker. We apply the proposed ITD-based speaker-grouped frequency fusion mechanism for different DOA estimation methods, namely the multiple signal classification, steered response power and a recently published method based on relative transfer function (RTF) vectors. In our experiments, we compare DOA estimation with different fusion mechanisms. For all considered DOA estimation methods, the proposed ITD-based speaker-grouped frequency fusion mechanism results in a higher DOA estimation accuracy compared with the narrowband and broadband fusion mechanisms. Daniel Fejgin, Elior Hadad, Sharon Gannot, Zbynek Koldovský, Simon Doclo |
ICASSP | 5 |
| 2024 | Microphone Subset Selection for the Weighted Prediction Error Algorithm Using a Group Sparsity PenaltyabstractReverberation can severely degrade the quality of speech signals recorded using microphones in an enclosure. In acoustic sensor networks with spatially distributed microphones, a similar dereverberation performance may be achieved using only a subset of all available microphones. Using the popular convex relaxation method, in this paper we propose to perform microphone subset selection for the weighted prediction error (WPE) multi-channel dereverberation algorithm by introducing a group sparsity penalty on the prediction filter coefficients. The resulting problem is shown to be solved efficiently using the accelerated proximal gradient algorithm. Experimental evaluation using measured impulse responses shows that the performance of the proposed method is close to the optimal performance obtained by exhaustive search, both for frequency-dependent as well as frequency-independent microphone subset selection. Furthermore, the performance using only a few microphones for frequency-independent microphone subset selection is only marginally worse than using all available microphones. Anselm Lohmann, Toon van Waterschoot, Jörg Bitzer, Simon Doclo |
ICASSP | 4 |
| 2024 | Multi-Microphone Noise Data Augmentation for DNN-Based Own Voice Reconstruction for Hearables in Noisy EnvironmentsabstractHearables with integrated microphones may offer communication benefits in noisy working environments, e.g. by transmitting the recorded own voice of the user. Systems aiming at reconstructing the clean and full-bandwidth own voice from noisy microphone recordings are often based on supervised learning. Recording a sufficient amount of noise required for training such a system is costly since noise transmission between outer and inner micro-phones varies individually. Previously proposed methods either do not consider noise, only consider noise at outer microphones or assume inner and outer microphone noise to be independent during training, and it is not yet clear whether individualized noise can benefit the training of and own voice reconstruction system. In this paper, we investigate several noise data augmentation techniques based on measured transfer functions to simulate multi-microphone noise. Using augmented noise, we train a multi-channel own voice reconstruction system. Experiments using real noise are carried out to investigate the generalization capability. Results show that incorporating augmented noise yields large benefits, in particular considering individualized noise augmentation leads to higher performance. Mattes Ohlenbusch, Christian Rollwage, Simon Doclo |
ICASSP | 3 |
| 2024 | Active Learning for Sound Event Classification Using Bayesian Neural Networks with Gaussian Variational PosteriorabstractManual annotation of audio material is cumbersome. Active learning aims at minimizing the annotation effort by iteratively selecting an acquisition batch of unlabeled data, asking a human to annotate the selected data and re-training a classifier until an annotation budget is depleted. In this paper we propose the Gaussian-dense active learning (GDAL) algorithm to train a sound event classifier. The classifier is a Bayesian neural network where the weights are normally distributed. This is in contrast to conventional neural networks where weights are not distributed, but have assigned values. The Bayesian nature of the classifier empowers GDAL to select acquisition batches from a set of unlabeled audio clips based on their estimated informativeness. Evaluation results on the UrbanSound8k dataset show that GDAL outperforms a state-of-the-art algorithm based on medoid active learning for all considered annotation budgets and an algorithm based on dropout active learning for sufficiently large annotation budgets. Stepan Shishkin, Danilo Hollosi, Stefan Goetze, Simon Doclo |
ICASSP | 4 |
| 2024 | Binaural Speech Enhancement Using Deep Complex Convolutional Transformer NetworksabstractStudies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an encoder-decoder architecture and a complex multi-head attention transformer. The model is trained to estimate individual complex ratio masks in the time-frequency domain for the left and right-ear channels of binaural hearing devices. The model is trained using a novel loss function that incorporates the preservation of spatial information along with speech intelligibility improvement and noise reduction. Simulation results for acoustic scenarios with a single target speaker and isotropic noise of various types show that the proposed method improves the estimated binaural speech intelligibility and preserves the binaural cues better in comparison with several baseline algorithms. Vikas Tokala, Eric Grinstein, Mike Brookes, Simon Doclo, Jesper Jensen 0001, Patrick A. Naylor |
ICASSP | 4 |
| 2024 | Effect of Target Signals and Delays on Spatially Selective Active Noise Control for Open-Fitting HearablesabstractSpatially selective active noise control (ANC) hearables are designed to reduce unwanted noise from certain directions while preserving desired sounds from other directions. In previous studies, the target signal has been defined either as the delayed desired component in one of the reference microphone signals or as the desired component in the error microphone signal without any delay. In this paper, we systematically investigate the influence of delays in different target signals on the ANC performance and provide an intuitive explanation for how the system obtains the desired signal. Simulations were conducted on a pair of open-fitting hearables for localized speech and noise sources in an anechoic environment. The performance was assessed in terms of noise reduction, signal quality and control effort. Results indicate that optimal performance is achieved without delays when the target signal is defined at the error microphone, whereas causality necessitates delays when the target signal is defined at the reference microphone. The optimal delay is found to be the acoustic delay between this reference microphone and the error microphone from the desired source. Tong Xiao 0008, Simon Doclo |
ICASSP | 2 |
| 2024 | Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
Marvin Tammen, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Shoko Araki, Simon Doclo |
INTERSPEECH | 6 |
| 2024 | Channel-Configurable Deep Wireless Speech TransmissionabstractThe proliferation of edge-based wireless speech applications necessitates the development of resource-efficient, low-latency speech communication systems capable of functioning across diverse communication channel conditions. Ensuring intelligible speech communication under conditions of constrained resources and low-latency presents a challenging problem within the domain of speech transmission. In this paper, we introduce a very low-latency configurable speech transmission system leveraging joint source-channel coding and deep neural networks (DNNs). Our proposed system is a unified deep neural network system engineered to operate effectively across a wide range of wireless communication channel scenarios. The system encompasses both a joint source-channel encoder and a joint source-channel decoder, each with access to channel state information (CSI). In this context, CSI signifies the type of fading in the wireless channel. Notably, our system has a total latency of 2 ms. Through extensive simulations, we empirically demonstrate that the proposed configurable system closely approximates the performance of ideal systems specifically tailored to individual wireless channel scenarios. Our evaluation is rooted in the assessment of instrumental measures of speech quality and intelligibility, affirming the efficacy of our system in diverse and resource-constrained communication contexts. Mohammad Bokaei, Jesper Jensen 0001, Simon Doclo, Jan Østergaard |
WCNC | 3 |
| 2024 | Speech-Aware Binaural DOA Estimation Utilizing Periodicity and Spatial Features in Convolutional Neural NetworksabstractIn recent years, several supervised learning-based approaches have been proposed for estimating the direction of arrival (DOA) of a single talker in noisy and reverberant environments. In the absence of auxiliary information, such as a voice activity detector (VAD), the estimated DOA may be erroneous due to speech pauses or noise dominance. In this paper, we consider a speech-aware DOA estimation system for binaural hearing aids, which does not require a separate VAD. This system utilizes a combination of spatial features with an auditory-inspired periodicity feature called periodicity degree (PD) as input features of a convolutional neural network (CNN). Using speech and non-speech signals during the training, the CNN can capture the harmonic structure encoded in the PD features, thereby distinguishing speech from non-speech portions and simultaneously mapping spatial features to sound source DOA upon speech detection. To investigate the benefit of using PD features for speech-aware DOA estimation, we evaluated the performance of speech-aware systems that utilized either broadband or narrowband feature combinations compared to baseline systems. We propose to use a novel narrowband feature combination consisting of the narrowband cross-power spectrum (CPS) as the spatial feature and a new subband-averaged representation of PD features. The broadband feature combination consisted of the generalized cross-correlation with phase transform (GCC-PHAT) and the broadband PD features. The baseline systems considered in this work consisted of a CNN that exploits only a spatial feature, cascaded with a VAD. Evaluations in reverberant environments with different background noises for both static and dynamic single-talker scenarios demonstrate that incorporating the PD feature in conjunction with any type of spatial feature provides an advantage for binaural DOA estimation in terms of accuracy and angular error. Reza Varzandeh, Simon Doclo, Volker Hohmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting A Calibrated External Microphone ArrayabstractRecently, a relative transfer function (RTF) vector-based method has been proposed to estimate the direction of arrival (DOA) of a target speaker for a binaural hearing aid setup, assuming the availability of external microphones. This method exploits the external microphones to estimate the RTF vector corresponding to the binaural hearing aid and constructs a one-dimensional spatial spectrum by comparing the estimated RTF vector against a database of anechoic prototype RTF vectors for several directions. In this paper, we assume the availability of a calibrated array of external microphones, which is characterized by a second database of anechoic prototype RTF vectors. We propose a method where the external microphones are not only exploited for RTF vector estimation but also assist in estimating the DOA of the target speaker. Based on the estimated RTF vector for all microphones and the prototype RTF databases of the binaural hearing aid and the external microphone array, a two-dimensional spatial spectrum is constructed from which the DOA is estimated. Experimental results for a reverberant environment with diffuse-like noise show that assisted DOA estimation outperforms DOA estimation where the prototype database characterizing the external microphone array is not used. Daniel Fejgin, Simon Doclo |
ICASSP | 2 |
| 2023 | Geometry-Aware DOA Estimation Using a Deep Neural Network with Mixed-Data Input FeaturesabstractUnlike model-based direction of arrival (DoA) estimation algorithms, supervised learning-based DoA estimation algorithms based on deep neural networks (DNNs) are usually trained for one specific microphone array geometry, resulting in poor performance when applied to a different array geometry. In this paper we illustrate the fundamental difference between supervised learning-based and model-based algorithms leading to this sensitivity. Aiming at designing a supervised learning-based DoA estimation algorithm that generalizes well to different array geometries, in this paper we propose a geometry-aware DoA estimation algorithm. The algorithm uses a fully connected DNN and takes mixed data as input features, namely the time lags maximizing the generalized cross-correlation with phase transform and the microphone coordinates, which are assumed to be known. Experimental results in a reverberant scenario demonstrate the flexibility of the proposed algorithm towards different array geometries and show that the proposed algorithm outperforms model-based algorithms such as steered response power with phase transform. Ulrik Kowalk, Simon Doclo, Jörg Bitzer |
ICASSP | 2 |
| 2023 | Dereverberation in Acoustic Sensor Networks Using weighted Prediction Error with Microphone-Dependent Prediction DelaysabstractIn the last decades several multi-microphone speech dereverberation algorithms have been proposed, among which the weighted prediction error (WPE) algorithm. In the WPE algorithm, a prediction delay is required to reduce the correlation between the prediction signals and the direct component in the reference microphone signal. In compact arrays with closely-spaced microphones, the prediction delay is often chosen microphone-independent. In acoustic sensor networks with spatially distributed microphones, large time-differences-of-arrival (TDOAs) of the speech source between the reference microphone and other microphones typically occur. Hence, when using a microphone-independent prediction delay1the reference and prediction signals may still be significantly correlated, leading to distortion in the dereverberated output signal. In order to decorrelate the signals, in this paper we propose to apply TDOA compensation with respect to the reference microphone, resulting in microphone-dependent prediction delays for the WPE algorithm. We consider both optimal TDOA compensation using crossband filtering in the short-time Fourier transform domain as well as band-to-band and integer delay approximations. Simulation results for different reverberation times using oracle as well as estimated TDOAs clearly show the benefit of using microphone-dependent prediction delays. Anselm Lohmann, Toon van Waterschoot, Jörg Bitzer, Simon Doclo |
ICASSP | 4 |
| 2023 | Joint Online Estimation of Early and Late Residual Echo PSD for Residual Echo SuppressionabstractIn hands-free telephony and other distant-talking applications, an acoustic echo cancellation system is typically required, where a short adaptive filter is often used in practice to achieve fast convergence at low computational cost. This may result in late residual echo (LRE) remaining due to under-modeling of the echo path and early residual echo (ERE) due to filter misalignment. Both residual echo components can be suppressed using a postfilter in the subband domain, which requires accurate estimates of the power spectral density (PSD) of the ERE and LRE components. State-of-the-art methods estimate the ERE and LRE PSDs independently of each other, where the ERE PSD is estimated by simply multiplying the loudspeaker PSD with a frequency-dependent scalar and the LRE PSD is estimated using a recursive estimator based on frequency-dependent reverberation scaling and decay parameters. In this paper, we propose to extend the ERE PSD estimator from a scalar to a moving average filter on the loudspeaker PSD. In addition, we propose a signal-based method to jointly estimate all model parameters for the ERE and LRE PSD estimators in online mode, and derive two gradient-descent-based algorithms to simultaneously update the model parameters by minimizing the mean squared log error. The proposed method is compared with state-of-the-art methods in terms of estimation accuracy of the model parameters as well as the residual echo PSDs. Simulation results using both artificially generated as well as measured impulse responses show that the proposed method outperforms state-of-the-art methods for all considered scenarios. Naveen Kumar Desiraju, Simon Doclo, Markus Buck, Tobias Wolff |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Parameter Estimation Procedures for Deep Multi-Frame MVDR Filtering for Single-Microphone Speech EnhancementabstractAiming at exploiting temporal correlations across consecutive time frames in the short-time Fourier transform (STFT) domain, multi-frame algorithms for single-microphone speech enhancement have been proposed, which apply a complex-valued filter to the noisy STFT coefficients. Typically, the multi-frame filter coefficients are either estimated directly using deep neural networks or a certain filter structure is imposed, e.g., the multi-frame minimum variance distortionless response (MFMVDR) filter structure. Recently, it was shown that integrating the fully differentiable MFMVDR filter into an end-to-end supervised learning framework employing temporal convolutional networks (TCNs) allows for a high estimation accuracy of the required parameters, i.e., the speech inter-frame correlation vector and the interference covariance matrix. In this paper, we investigate different covariance matrix structures, namely Hermitian positive-definite, Hermitian positive-definite Toeplitz, and rank-1. The main differences between the considered matrix structures lie in the number of parameters that need to be estimated by the TCNs as well as the required linear algebra operations, yielding a different computational complexity. For example, when assuming a rank-1 matrix structure, we show that the MFMVDR filter can be written as a linear combination of the TCN outputs, significantly reducing computational complexity. In addition, we consider a covariance matrix estimation procedure based on recursive smoothing, where the smoothing factors are estimated using TCNs. Experimental results on the deep noise suppression challenge dataset show that the estimation procedure using the Hermitian positive-definite matrix structure yields the best performance, closely followed by the rank-1 matrix structure at a much lower complexity. Furthermore, it is shown for the best-performing MFMVDR filters that imposing the MFMVDR filter structure instead of directly estimating the multi-frame filter coefficients slightly but consistently improves the speech enhancement performance. Marvin Tammen, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Optimization of a Fixed Virtual Sensing Feedback ANC Controller For In-Ear Headphones with Multiple LoudspeakersabstractIn this paper we consider an in-ear headphone equipped with an inner microphone and multiple loudspeakers and we propose an optimization procedure with a convex objective function to derive a fixed multi-loudspeaker ANC controller aiming at minimizing the sound pressure at the ear drum. Based on the virtual microphone arrangement (VMA) technique and measured acoustic paths between the loudspeakers and the ear drum, the FIR filters of the ANC controller are jointly optimized to minimize the power spectral density at the ear drum, subject to design and stability constraints. For an in-ear headphone with two loudspeakers, the proposed multi-loudspeaker VMA controller is compared to two single-loudspeaker VMA controllers. Simulation results with diffuse noise show that the multi-loudspeaker VMA controller effectively improves the attenuation by up to about 10 dB for frequencies below 300 Hz when compared to both single-loudspeaker VMA controllers. Piero Rivera Benois, Reinhild Roden, Matthias Blau, Simon Doclo |
ICASSP | 4 |
| 2022 | A Class of Pareto Optimal Binaural BeamformersabstractThe objective of binaural multi-microphone speech enhancement algorithms can be viewed as a multi-criteria design problem as there are several requirements to be met. The objective is not only to extract the target speaker without distortion, but also to suppress interfering sources (e.g., competing speakers) and ambient background noise, while preserving the auditory impression of the complete acoustic scene. Such a multi-objective problem (MOP) can be solved using a Pareto frontier, which provides a useful trade-off between the different criteria. In this paper, we propose a unified Pareto optimization framework, which is achieved by defining a generalized mean squared error (MSE) cost function, derived from a MOP. The solution to the multi-criteria problem is grounded on a solid mathematical foundation. The MSE cost function consists of a weighted sum of speech distortion (SD), partial interference reduction (IR), and partial noise reduction (NR) terms with scaling parameters that control the amount of IR and NR. The filter minimizing this generalized cost function, denoted Pareto optimal binaural multichannel Wiener filter (Pareto-BMWF), constitutes a generalization of various binaural MWF-based and binaural MVDR-based beamformers. This solution is optimal for any set of parameters. The improved speech enhancement capabilities are experimentally demonstrated using real-signal recordings when estimation errors are present and the binaural cue preservation capabilities are analyzed. Elior Hadad, Simon Doclo, Sven Nordholm, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Deep Multi-Frame MVDR Filtering for Single-Microphone Speech EnhancementabstractMulti-frame algorithms for single-microphone speech enhancement, e.g., the multi-frame minimum variance distortionless response (MFMVDR) filter, are able to exploit speech correlation across adjacent time frames in the short-time Fourier transform (STFT) domain. Provided that accurate estimates of the required speech interframe correlation vector and the noise correlation matrix are available, it has been shown that the MFMVDR filter yields a substantial noise reduction while hardly introducing any speech distortion. Aiming at merging the speech enhancement potential of the MFMVDR filter and the estimation capability of temporal convolutional networks (TCNs), in this paper we propose to embed the MFMVDR filter within a deep learning framework. The TCNs are trained to map the noisy speech STFT coefficients to the required quantities by minimizing the scale-invariant signal-to-distortion ratio loss function at the MFMVDR filter output. Experimental results show that the proposed deep MFMVDR filter achieves a competitive speech enhancement performance on the Deep Noise Suppression Challenge dataset. In particular, the results show that estimating the parameters of an MFMVDR filter yields a higher performance in terms of PESQ and STOI than directly estimating the multi-frame filter or single-frame masks and than Conv-TasNet. Marvin Tammen, Simon Doclo |
ICASSP | 2 |
| 2021 | Robust Constrained MFMVDR Filters for Single-Channel Speech Enhancement Based on Spherical Uncertainty SetabstractAiming at exploiting speech correlation across consecutive time-frames in the short-time Fourier transform domain, the multi-frame minimum variance distortionless response (MFMVDR) filter for single-channel speech enhancement has been proposed. The MFMVDR filter requires an accurate estimate of the normalized speech correlation vector in order to avoid speech distortion and artifacts. In this paper we investigate the potential of using robust MVDR filtering techniques to estimate the normalized speech correlation vector as the vector maximizing the total signal output power within a spherical uncertainty set, which corresponds to imposing a quadratic inequality constraint. Whereas the singly-constrained (SC) MFMVDR filter only considers the quadratic inequality constraint to estimate the (non-normalized) speech correlation vector, the doubly-constrained (DC) MFMVDR filter integrates a linear normalization constraint into the optimization problem to directly estimate the normalized speech correlation vector. To set the upper bound of the quadratic inequality constraint for each time-frequency point, we propose to use a trained non-linear mapping function that depends on the a-priori signal-to-noise ratio (SNR). Experimental results for different speech signals, noise types and SNRs show that the proposed constrained approaches yield a more accurate estimate of the normalized speech correlation vector than a state-of-the-art maximum-likelihood (ML) estimator. An instrumental and a perceptual evaluation show that both constrained MFMVDR filters lead to less speech and noise distortion but a lower noise reduction than the ML-MFMVDR filter, where the DC-MFMVDR filter is preferred in terms of overall quality compared to the SC-MFMVDR and ML-MFMVDR filters. Dörte Fischer, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Performance Analysis of the Extended Binaural MVDR Beamformer With Partial Noise EstimationabstractBesides reducing undesired noise sources and limiting speech distortion, another important objective of a binaural noise reduction algorithm is the preservation of the binaural cues of all sound sources in the acoustic scene. In this paper, we consider the binaural minimum variance distortionless response beamformer with partial noise estimation (BMVDR-N), which allows to trade off between noise reduction performance and binaural cue preservation of the noise component by mixing the output signals of the BMVDR beamformer with the noisy reference microphone signals. For a directional noise source, it has been shown that incorporating an external microphone in addition to the head-mounted microphones enables both the noise reduction performance as well as the interaural time and level difference cues of the noise component to be improved in the output signals. In this paper, we consider an arbitrary noise field and analytically show that incorporating an external microphone in the BMVDR-N beamformer enables 1) a larger output signal-to-noise ratio (SNR) for the same mixing parameter, 2) the same output SNR for a larger mixing parameter, and 3) the same desired output magnitude squared coherence (MSC) of the noise component for a smaller mixing parameter to be obtained. The derived analytical expressions are firstly validated using simulated anechoic acoustic transfer functions, where the listener's head is modelled as a rigid sphere. Experimental results using recorded signals for a binaural hearing device setup in a reverberant environment also show that in a realistic scenario incorporating an external microphone in the BMVDR-N beamformer significantly improves the output SNR and reduces the mixing parameter that is required to obtain a desired output MSC of the noise component compared to using only the head-mounted microphones. Nico Gößling, Daniel Marquardt, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Improving Auditory Attention Decoding Performance of Linear and Non-Linear Methods using State-Space ModelabstractIdentifying the target speaker in hearing aid applications is crucial to improve speech understanding. Recent advances in electroencephalography (EEG) have shown that it is possible to identify the target speaker from single-trial EEG recordings using auditory attention decoding (AAD) methods. AAD methods reconstruct the attended speech envelope from EEG recordings, based on a linear least-squares cost function or non-linear neural networks, and then directly compare the reconstructed envelope with the speech envelopes of speakers to identify the attended speaker using Pearson correlation coefficients. Since these correlation coefficients are highly fluctuating, for a reliable decoding a large correlation window is used, which causes a large processing delay. In this paper, we investigate a state-space model using correlation coefficients obtained with a small correlation window to improve the decoding performance of the linear and the non-linear AAD methods. The experimental results show that the state-space model significantly improves the decoding performance. Ali Aroudi, Tobias de Taillez, Simon Doclo |
ICASSP | 3 |
| 2020 | Subspace-Based Speech Correlation Vector Estimation for Single-Microphone Multi-Frame MVDR FilteringabstractAiming at exploiting the speech correlation across consecutive timeframes in the short-time Fourier transform domain, the multi-frame minimum variance distortionless response (MFMVDR) filter for single-microphone speech enhancement has been proposed. This filter is designed to avoid speech distortion while minimizing the total signal output power. To compute the MFMVDR filter, an estimate of the highly time-varying normalized speech correlation vector is required. In this paper, we propose a subspace-based estimator for the normalized speech correlation vector based on the Q largest eigenvalues and their corresponding eigenvectors of the prewhitened noisy speech correlation matrix. Experimental results for different speech signals, noise types and signal-to-noise ratios show that the proposed subspace-based estimator yields the best results in terms of speech quality and noise reduction compared to a state-of-the-art maximum-likelihood estimator. Dörte Fischer, Simon Doclo |
ICASSP | 2 |
| 2020 | DNN-Based Speech Presence Probability Estimation for Multi-Frame Single-Microphone Speech EnhancementabstractMulti-frame approaches for single-microphone speech enhancement, e.g., the multi-frame minimum-power-distortionless-response (MFMPDR) filter, are able to exploit speech correlations across neighboring time frames. In contrast to single-frame approaches such as the Wiener gain, it has been shown that multi-frame approaches achieve a substantial noise reduction with hardly any speech distortion, provided that an accurate estimate of the correlation matrices and especially the speech interframe correlation (IFC) vector is available. Typical estimation procedures of the IFC vector require an estimate of the speech presence probability (SPP) in each time-frequency (TF) bin. In this paper, we propose to use a bi-directional long short-term memory deep neural network (DNN) to estimate the SPP for each TF bin. Aiming at achieving a robust performance, the DNN is trained for various noise types and within a large signal-to-noise-ratio range. Experimental results show that the MFMPDR in combination with the proposed data-driven SPP estimator yields an increased speech quality compared to a state-of-the-art model-based SPP estimator. Furthermore, it is confirmed that exploiting interframe correlations in the MFMPDR is beneficial when compared to the Wiener gain especially in adverse scenarios. Marvin Tammen, Dörte Fischer, Bernd T. Meyer, Simon Doclo |
ICASSP | 4 |
| 2020 | Exploiting Periodicity Features for Joint Detection and DOA Estimation of Speech Sources Using Convolutional Neural NetworksabstractWhile many algorithms deal with direction of arrival (DOA) estimation and voice activity detection (VAD) as two separate tasks, only a small number of data-driven methods have addressed these two tasks jointly. In this paper, a multi-input single-output convolutional neural network (CNN) is proposed which exploits a novel feature combination for joint DOA estimation and VAD in the context of binaural hearing aids. In addition to the well-known generalized cross correlation with phase transform (GCC-PHAT) feature, the network uses an auditory-inspired feature called periodicity degree (PD), which provides a broadband representation of the periodic structure of the signal. The proposed CNN has been trained in a multi-conditional training scheme across different signal-to-noise ratios. Experimental results for a single-talker scenario in reverberant environments show that by exploiting the PD feature, the proposed CNN is able to distinguish speech from non-speech signal blocks, thereby outperforming the baseline CNN in terms of DOA estimation accuracy. In addition, the results show that the proposed method is able to adapt to different unseen acoustic conditions and background noises. Reza Varzandeh, Kamil Adiloglu, Simon Doclo, Volker Hohmann |
ICASSP | 3 |
| 2020 | Adaptive Compressive Onset-Enhancement for Improved Speech Intelligibility in Noise and ReverberationabstractNear-end listening enhancement (NELE) algorithms aim to pre-process speech prior to playback via loudspeakers so as to maintain high speech intelligibility even when listening conditions are not optimal, e.g., due to noise or reverberation. Often NELE algorithms are designed for scenarios considering either only the detrimental effect of noise or only reverberation, but not both disturbances. In many typical applications scenarios, however, both factors are present. In this paper, we evaluate a new combination of a noise-dependent and a reverberation-dependent algorithm implemented in a common framework. Specifically, we use instrumental measures as well as subjective ratings of listening effort for acoustic scenarios with different reverberation times and realistic signal-to-noise ratios. The results show that the noise-dependent algorithm also performs well in reverberation, and that the combination of both algorithms can yield slightly better performance than the individual algorithms alone. This benefit appears to depend strongly on the specific acoustic condition, indicating that further work is required to optimize the adaptive algorithm behavior. Felicitas Bederna, Henning F. Schepker, Christian Rollwage, Simon Doclo, Arne Pusch, Jörg Bitzer, Jan Rennies |
INTERSPEECH | 4 |
| 2020 | Cognitive-Driven Binaural Beamforming Using EEG-Based Auditory Attention DecodingabstractIdentifying the target speaker in hearing aid applications is an essential ingredient to improve speech intelligibility. Recently, a least-squares-based auditory attention decoding (AAD) method has been proposed to identify the target speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers. Aiming at enhancing the target speaker and suppressing the interfering speaker and ambient noise, in this article, we propose a cognitive-driven speech enhancement system, consisting of a binaural beamformer which is steered based on AAD and estimated relative transfer function (RTF) vectors, which require estimates of the direction-of-arrivals (DOAs) of both speakers. For binaural beamforming and to generate reference signals for AAD, we consider either minimum-variance-distortionless-response (MVDR) beamformers or linearly-constrained-minimum-variance (LCMV) beamformers. Contrary to the binaural MVDR beamformer, the binaural LCMV beamformer allows to preserve the spatial impression of the acoustic scene and to control the suppression of the interfering speaker, which is important when intending to switch attention between speakers. The speech enhancement performance of the proposed system is evaluated in terms of the binaural signal-to-interference-plus-noise ratio (SINR) improvement in anechoic and reverberant conditions. Furthermore, we investigate the impact of RTF and DOA estimation errors and AAD errors on the speech enhancement performance. The experimental results show that the proposed system using LCMV beamformers yields a larger decoding performance and binaural SINR improvement compared to using MVDR beamformers. Ali Aroudi, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Online Estimation of Reverberation Parameters For Late Residual Echo SuppressionabstractIn hands-free telephony and other distant-talk applications, often a short AEC filter is used to achieve fast convergence at low computational cost. As a result, a significant amount of late residual echo (LRE) may remain, especially in highly reverberant environments. This LRE can be suppressed using a postfilter in the subband domain, which requires an estimate of the power spectral density (PSD) of the LRE. To estimate the LRE PSD, an exponentially decaying model with frequency-dependent reverberation scaling and decay parameters has frequently been assumed. State-of-the-art methods estimate both reverberation parameters independently of each other, either in offline or in online mode. In this article, we propose two signal-based methods (i.e. output error and equation error) to jointly estimate both reverberation parameters in online mode. The estimated parameters are then used to generate an estimate for the LRE PSD, which is fed into a postfilter for the purpose of late residual echo suppression. We derive several gradient-descent-based algorithms to simultaneously update both reverberation parameters, minimizing either the mean squared error or the mean squared log error cost function. The proposed methods are compared with state-of-the-art methods in terms of the accuracy of the estimated reverberation parameters and the corresponding LRE PSD estimate. Extensive simulation results using both artificial as well as measured room impulse responses show that the proposed output error method with mean squared log error minimization outperforms state-of-the-art methods in all considered scenarios. Naveen Kumar Desiraju, Simon Doclo, Markus Buck, Tobias Wolff |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Integrated Sidelobe Cancellation and Linear Prediction Kalman Filter for Joint Multi-Microphone Speech Dereverberation, Interfering Speech Cancellation, and Noise ReductionabstractIn multi-microphone speech enhancement, reverberation as well as additive noise and/or interfering speech are commonly suppressed by deconvolution and spatial filtering, e.g., using multi-channel linear prediction (MCLP) on the one hand and beamforming, e.g., a generalized sidelobe canceler (GSC), on the other hand. In this article, we consider several reverberant speech components, whereof some are to be dereverberated and others to be canceled, as well as a diffuse (e.g., babble) noise component to be suppressed. In order to perform both deconvolution and spatial filtering, we integrate MCLP and the GSC into a novel architecture referred to as integrated sidelobe cancellation and linear prediction (ISCLP), where the sidelobe-cancellation (SC) filter and the linear prediction (LP) filter operate in parallel, but on different microphone signal frames. Within ISCLP, we estimate both filters jointly by means of a single Kalman filter. We further propose a spectral Wiener gain post-processor, which is shown to relate to the Kalman filter's posterior state estimate. The presented ISCLP Kalman filter is benchmarked against two state-of-the-art approaches, namely first a pair of alternating Kalman filters respectively performing dereverberation and noise reduction, and second an MCLP+GSC Kalman filter cascade. While the ISCLP Kalman filter is roughly $M^2$ times less expensive than both reference algorithms, where $M$ denotes the number of microphones, it is shown to perform at least similarly as compared to the former, and to outperform the latter. A MATLAB implementation is available. Thomas Dietzen, Simon Doclo, Marc Moonen, Toon van Waterschoot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Square Root-Based Multi-Source Early PSD Estimation and Recursive RETF Update in Reverberant Environments by Means of the Orthogonal Procrustes ProblemabstractMulti-channel short-time Fourier transform (STFT) domain-based processing of reverberant microphone signals commonly relies on power-spectral-density (PSD) estimates of early source images, where early refers to reflections contained within the same STFT frame. State-of-the-art approaches to multi-source early PSD estimation, given an estimate of the associated relative early transfer functions (RETFs), conventionally minimize the approximation error defined with respect to the early correlation matrix, requiring non-negative inequality constraints on the PSDs. Instead, we here propose to factorize the early correlation matrix and minimize the approximation error defined with respect to the early-correlation-matrix square root. The proposed minimization problem-constituting a generalization of the so-called orthogonal Procrustes problem-seeks a unitary matrix and the square roots of the early PSDs up to an arbitrary complex argument, whereby non-negative inequality constraints become redundant. A solution is obtained iteratively, requiring one singular value decomposition (SVD) per iteration. The estimated unitary matrix and early PSD square roots further allow to recursively update the RETF estimate, which is not inherently possible in the conventional approach. An estimate of the said early-correlation-matrix square root itself is obtained by means of the generalized eigenvalue decomposition (GEVD), where we further propose to restore non-stationarities by desmoothing the generalized eigenvalues in order to compensate for inevitable recursive averaging. Simulation results indicate fast convergence of the proposed multi-source early PSD estimation approach in only one iteration if initialized appropriately, and better performance as compared to the conventional approach. A MATLAB implementation is available. Thomas Dietzen, Simon Doclo, Marc Moonen, Toon van Waterschoot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Binaural LCMV Beamforming With Partial Noise EstimationabstractBesides reducing undesired sources, i.e., interfering sources and background noise, another important objective of a binaural beamforming algorithm is to preserve the spatial impression of the acoustic scene, which can be achieved by preserving the binaural cues of all sound sources. While the binaural minimum variance distortionless response (BMVDR) beamformer provides a good noise reduction performance and preserves the binaural cues of the desired source, it does not allow to control the reduction of the interfering sources and distorts the binaural cues of the interfering sources and the background noise. Hence, several extensions have been proposed. First, the binaural linearly constrained minimum variance (BLCMV) beamformer uses additional constraints, enabling to control the reduction of the interfering sources while preserving their binaural cues. Second, the BMVDR with partial noise estimation (BMVDR-N) mixes the output signals of the BMVDR with the noisy reference microphone signals, enabling to control the binaural cues of the background noise. Aiming at merging the advantages of both extensions, in this paper we propose the BLCMV with partial noise estimation (BLCMV-N). We show that the output signals of the BLCMV-N can be interpreted as a mixture between the noisy reference microphone signals and the output signals of a BLCMV using an adjusted interference scaling parameter. We provide a theoretical comparison between the BMVDR, the BLCMV, the BMVDR-N and the proposed BLCMV-N in terms of noise and interference reduction performance and binaural cue preservation. Experimental results using recorded signals as well as the results of a perceptual listening test show that the BLCMV-N is able to preserve the binaural cues of an interfering source (like the BLCMV), while enabling to trade off between noise reduction performance and binaural cue preservation of the background noise (like the BMVDR-N). Nico Gößling, Elior Hadad, Sharon Gannot, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Acoustic Feedback Suppression for Multi-Microphone Hearing Devices Using a Soft-Constrained Null-Steering BeamformerabstractAcoustic feedback occurs in hearing aids due to the coupling between the hearing aid loudspeaker and microphone(s). In order to reduce the acoustic feedback, adaptive filters are commonly used to estimate the feedback contribution in the microphone(s). While theoretically allowing for perfect feedback cancellation, in practice the adaptive filter typically converges to a biased optimal solution due to the closed-loop acoustical system of the hearing aid. Previously it has therefore been proposed to suppress the acoustic feedback contribution for an earpiece with multiple integrated microphones and loudspeakers using a fixed null-steering beamformer and hence avoiding a biased adaption. While previous null-steering beamforming approaches aimed at perfect preservation of the incoming signal using its relative transfer function (RTF), in this article we propose to use a soft constraint that allows to trade off between incoming signal preservation and feedback suppression. We formulate the computation of the beamformer coefficients both as a least-squares optimization procedure, aiming to minimize the residual feedback power, and as a min-max optimization procedure, aiming to directly maximize the maximum stable gain of the hearing aid. Experimental evaluations were performed using measured acoustic feedback paths from a custom earpiece with two microphones in the vent and a third microphone in the concha. Results show that the proposed fixed null-steering beamformer using the RTF-based soft constraint provides a reduction of the acoustic feedback by 7-8 dB compared to the previously proposed RTF-based hard constraint while limiting the distortions of the incoming signal in the beamformer output. Henning F. Schepker, Sven Nordholm, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Cognitive-driven Binaural LCMV Beamformer Using EEG-based Auditory Attention DecodingabstractIdentifying the target speaker in hearing aid applications is an essential ingredient to improve speech intelligibility. To identify the target speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers, an auditory attention decoding (AAD) method was recently proposed. Aiming at enhancing the target speaker and suppressing the interfering speaker and ambient noise, in this paper we propose a cognitive-driven speech enhancement system, consisting of a direction-of-arrival (DOA) estimator, steerable beamformers and AAD. To preserve the spatial impression of the acoustic scene, which is important when intending to switch attention between speakers, the proposed system only partially suppresses the interfering speaker. The speech enhancement performance of the proposed system is evaluated in terms of the signal-to-interference-plus-noise ratio (SINR) improvement in anechoic and reverberant conditions. The experimental results show that the proposed system can obtain a considerably large SINR improvement (between 3.1 dB and 7.5 dB) in both conditions. Ali Aroudi, Simon Doclo |
ICASSP | 2 |
| 2019 | RTF-steered Binaural MVDR Beamforming Incorporating an External Microphone for Dynamic Acoustic ScenariosabstractA well-known binaural noise reduction algorithm is the binaural minimum variance distortionless response beamformer, which can be steered using the relative transfer function (RTF) vectors of the desired source. In this paper, we consider the recently proposed spatial coherence (SC) method to estimate the RTF vectors, requiring an additional external microphone that is spatially separated from the head-mounted microphones. Although the SC method provides a biased estimate of the RTF between the head-mounted microphones and the external microphone, we show that this bias is real-valued and only depends on the SNR in the external microphone. We propose to use the SC method to estimate the extended RTF vectors that also incorporate the external microphone, enabling to filter the external microphone signal in conjunction with the head-mounted microphones. Evaluation results using recorded signals of a moving speaker in diffuse noise show that the SC method yields a slightly better performance than the widely used covariance whitening method at a much lower computational complexity. Nico Gößling, Simon Doclo |
ICASSP | 2 |
| 2019 | Joint Estimation of RETF Vector and Power Spectral Densities for Speech Enhancement Based on Alternating Least SquaresabstractThe multi-channel Wiener filter (MWF) is a well-known multi-microphone speech enhancement technique, aiming at improving the quality of the recorded speech signals in noisy and reverberant environments. Assuming that reverberation and ambient noise can be modeled as a diffuse sound field and the spatial coherence of the residual noise is known, the MWF requires estimates of the relative early transfer function (RETF) vector of the target speaker as well as the power spectral densities (PSDs) of the target, diffuse and residual noise component. RETF vector and PSD estimation is often decoupled, where one quantity is estimated independently of the other quantity. In this paper, we propose to jointly estimate the RETF vector and all PSDs by minimizing the Frobenius norm of a model-based error matrix using an alternating least squares method. Experimental results using different dynamic acoustic scenarios with a moving speaker show that the proposed method leads to a larger MWF performance than a state-of-the-art method based on covariance whitening. Marvin Tammen, Simon Doclo, Ina Kodrasi |
ICASSP | 2 |
| 2019 | Non-Intrusive Speech Quality Prediction Using Modulation Energies and LSTM-NetworkabstractMany signal processing algorithms have been proposed to improve the quality of speech recorded in the presence of noise and reverberation. Perceptual measures, i.e., listening tests, are usually considered the most reliable way to evaluate the quality of speech processed by such algorithms but are costly and time-consuming. Consequently, speech enhancement algorithms are often evaluated using signal-based measures, which can be either intrusive or non-intrusive. As the computation of intrusive measures requires a reference signal, only non-intrusive measures can be used in applications for which the clean speech signal is not available. However, many existing non-intrusive measures correlate poorly with the perceived speech quality, particularly when applied over a wide range of algorithms or acoustic conditions. In this paper, we propose a novel non-intrusive measure of the quality of processed speech that combines modulation energy features and a recurrent neural network using long short-term memory cells. We collected a dataset of perceptually evaluated signals representing several acoustic conditions and algorithms and used this dataset to train and evaluate the proposed measure. Results show that the proposed measure yields higher correlation with perceptual speech quality than that of benchmark intrusive and non-intrusive measures when considering various categories of algorithms. Although the proposed measure is sensitive to mismatch between training and testing, results show that it is a useful approach to evaluate specific algorithms over a wide range of acoustic conditions and may, thus, become particularly useful for real-time selection of speech enhancement algorithm settings. Benjamin Cauchi, Kai Siedenburg, João Felipe Santos, Tiago H. Falk, Simon Doclo, Stefan Goetze |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Comparative Analysis of Generalized Sidelobe Cancellation and Multi-Channel Linear Prediction for Speech Dereverberation and Noise ReductionabstractFor blind speech dereverberation, two frameworks are commonly used: on the one hand, the multi-channel linear prediction (MCLP) framework, and on the other hand, data-dependent beamforming, e.g., the generalized sidelobe canceler (GSC) framework. The MCLP framework is designed to perform deconvolution and hence has gained increased prominence in blind speech dereverberation. The GSC framework is commonly used for noise reduction, but may be applied for dereverberation as well. In previous work, we have shown that for the noiseless case, MCLP and the GSC yield in theory mathematically equivalent results in terms of dereverberation. In this paper, we assume additional coherent as well as incoherent-noise components and formally analyze and compare both frameworks in terms of dereverberation and noise reduction performance. Both the theoretical analysis and time domain simulation results demonstrate that unlike the GSC, MCLP expectably shows limited performance in terms of noise reduction, while both perform equally well in terms of dereverberation, provided that the GSC blocking matrix achieves complete blocking of the early reverberant-speech component and sufficiently many microphones are available. In case of incomplete blocking, however, the GSC performs inferior to MCLP in terms of dereverberation, as shown in short-time Fourier transform domain simulations. Thomas Dietzen, Ann Spriet, Wouter Tirry, Simon Doclo, Marc Moonen, Toon van Waterschoot |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Null-Steering Beamformer-Based Feedback Cancellation for Multi-Microphone Hearing Aids With Incoming Signal PreservationabstractIn hearing aids, acoustic feedback occurs due to the coupling between the hearing aid loudspeaker and microphone(s). In order to reduce the acoustic feedback, adaptive filters are commonly used to estimate the feedback contribution in the microphone(s). While theoretically allowing for perfect feedback cancellation, in practice the adaptive filter converges to an optimal solution that is typically biased due to the closed-loop acoustical system of the hearing aid. In order to avoid the adaptation to a biased optimal solution, in this paper we propose to use a fixed beamformer to cancel the acoustic feedback contribution for an earpiece with multiple integrated microphones and loudspeakers. By steering a spatial null in the direction of the hearing aid loudspeaker, we show that theoretically perfect feedback cancellation can be achieved. While previous null-steering beamforming approaches did not control for distortions of the incoming signal, in this paper we propose to incorporate a constraint based on the relative transfer function (RTF) of the incoming signal, aiming to perfectly preserve this signal. We formulate the computation of the beamformer coefficients both as a least-squares optimization procedure, aiming to minimize the residual feedback power, and as a min-max optimization procedure, aiming to directly maximize the maximum stable gain of the hearing aid. Experimental results using measured acoustic feedback paths from a custom earpiece with two microphones in the vent and a third microphone in the concha show that the proposed fixed null-steering beamformer using the RTF-based constraint provides a reduction of the acoustic feedback and substantially increases the added stable gain while preserving the incoming signal. This can even be achieved for unknown acoustic feedback paths and incoming signal directions. Henning F. Schepker, Sven Nordholm, Linh Thi Thuc Tran, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | EEG-Based Auditory Attention Decoding Using Steerable Binaural Superdirective BeamformerabstractDuring the last decades significant progress in multi-microphone speech enhancement algorithms has been made for hearing aids. However, the performance of many algorithms depends on identifying the target speaker to be enhanced. To identify the target speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers, an auditory attention decoding (AAD) method was recently proposed. This AAD method however requires the clean speech signals of both the attended and the unattended speaker as reference signals for decoding. Since in practice only microphone signals, containing several undesired acoustic components, are available, in this paper we explore the potential of using steerable binaural superdirective beamformer for generating appropriate reference signals for decoding. The experimental results show that using steerable superdirective beamformer output signals improves the decoding performance compared to using the noisy microphone signals as reference signals. Ali Aroudi, Daniel Marquardt, Simon Doclo |
ICASSP | 3 |
| 2018 | Joint Late Reverberation and Noise Power Spectral Density Estimation in a Spatially Homogeneous Noise FieldabstractMany multi-channel dereverberation and noise reduction techniques such as the multi-channel Wiener filter (MWF) require an estimate of the late reverberation and noise power spectral densities (PSDs). State-of-the-art multi-channel methods for estimating the late reverberation PSD typically assume that the noise PSD matrix is known. Instead of assuming that the noise PSD matrix is known, in this paper we model the noise as a spatially homogeneous sound field with an unknown time-varying PSD and a known time-invariant spatial coherence matrix. Based on this model, two joint estimators of the late reverberation and noise PSDs are proposed, i.e., a non-blocking-based estimator which simultaneously estimates the target signal, late reverberation, and noise PSDs, and a blocking-based estimator which first estimates the late reverberation and noise PSDs at the output of a blocking matrix aiming to block the target signal. Experimental results show that the proposed blocking-based estimator yields the best performance when used in an MWF, even resulting in a similar or better performance than a state-of-the-art blocking-based estimator of the late reverberation PSD which assumes that the noise PSD matrix is known. Ina Kodrasi, Simon Doclo |
ICASSP | 2 |
| 2018 | Complexity Reduction of Eigenvalue Decomposition-Based Diffuse Power Spectral Density Estimators Using the Power MethodabstractIn noisy and reverberant environments speech enhancement techniques such as the multi-channel Wiener filter (MWF) can be used to improve speech quality and intelligibility. Assuming that reverberation and ambient noise can be modeled as diffuse sound fields, such techniques require an estimate of the diffuse power spectral density (PSD). Recently a multi-channel diffuse PSD estimator based on the eigenvalue decomposition (EVD) of the prewhitened signal PSD matrix was proposed. The EVD-based PSD estimator is advantageous in comparison to other state-of-the-art PSD estimators, since it does not require knowledge of the relative early transfer functions of the target signal. However, computing the EVD can be computationally expensive, particularly when the number of microphones is large. In this paper we propose to reduce the complexity of the EVD-based PSD estimator by using the iterative power method to compute the eigenvalues. Since the EVD-based PSD estimator only requires the largest eigenvalues, the full EVD is not required and the power method is a well suited computationally efficient technique to estimate these eigenvalues. Experimental results show that using the PSD estimated via the power method in an MWF yields a very similar performance as using the PSD estimated via the full EVD. Marvin Tammen, Ina Kodrasi, Simon Doclo |
ICASSP | 3 |
| 2018 | Evaluation and Comparison of Late Reverberation Power Spectral Density EstimatorsabstractReduction of late reverberation can be achieved using spatio-spectral filters, such as the multichannel Wiener filter. To compute this filter, an estimate of the late reverberation power spectral density (PSD) is required. In recent years, a multitude of late reverberation PSD estimators have been proposed. In this paper, these estimators are categorized into several classes, their relations and differences are discussed, and a comprehensive experimental comparison is provided. To compare their performance, simulations in controlled as well as practical scenarios are conducted. It is shown that a common weakness of spatial coherence-based estimators is their performance in high direct-to-diffuse ratio conditions. To mitigate this problem, a correction method is proposed and evaluated. It is shown that the proposed correction method can decrease the speech distortion without significantly affecting the reverberation reduction. Sebastian Braun, Adam Kuklasinski, Ofer Schwartz, Oliver Thiergart, Emanuël A. P. Habets, Sharon Gannot, Simon Doclo, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2018 | Analysis of Eigenvalue Decomposition-Based Late Reverberation Power Spectral Density EstimationabstractMany speech dereverberation techniques require an estimate of the late reverberation power spectral density (PSD). State-of-the-art multichannel methods for estimating the late reverberation PSD typically rely on first, an estimate of the relative transfer functions (RTFs) of the target signal; second, a model for the spatial coherence matrix of the late reverberation; and finally, an estimate of the reverberant speech or reverberant and noisy speech PSD matrix. The RTFs, the spatial coherence matrix, and the speech PSD matrix are all prone to modeling and estimation errors in practice, with the RTFs being particularly difficult to estimate accurately, especially in highly reverberant and noisy scenarios. Recently, we proposed an eigenvalue decomposition (EVD)-based late reverberation PSD estimator, which does not require an estimate of the RTFs. In this paper, this EVD-based PSD estimator is further analyzed and its estimation accuracy and computational complexity are analytically compared to a state-of-the-art maximum likelihood (ML) based PSD estimator. It is shown that for perfect knowledge of the RTFs, spatial coherence matrix, and reverberant speech PSD matrix, the ML-based and the EVD-based PSD estimates are both equal to the true late reverberation PSD. In addition, it is shown that for erroneous RTFs but perfect knowledge of the spatial coherence matrix and reverberant speech PSD matrix, the ML-based PSD estimate is larger than or equal to the true late reverberation PSD, whereas the EVD-based PSD estimate is obviously still equal to the true late reverberation PSD. Finally, it is shown that when modeling and estimation errors occur in all quantities, the ML-based PSD estimate is larger than or equal to the EVD-based PSD estimate. Simulation results for several realistic acoustic scenarios demonstrate the advantages of using the EVD-based PSD estimator in a multichannel Wiener filter, yielding a significantly better performance than the ML-based PSD estimator. Ina Kodrasi, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Interaural Coherence Preservation for Binaural Noise Reduction Using Partial Noise Estimation and Spectral PostfilteringabstractThe objective of binaural speech enhancement algorithms is to reduce the undesired noise component, while preserving the desired speech source and the binaural cues of all sound sources. For the scenario of a single desired speech source in a diffuse noise field, an extension of the binaural multichannel Wiener filter (MWF), namely the MWF-IC, has been recently proposed, which aims to preserve the interaural coherence (IC) of the noise component. However, due to the large complexity of the MWF-IC, in this paper we propose several alternative algorithms at a lower computational complexity. First, we consider a quasi-distortionless version of the MWF-IC, denoted as minimum-variance-distortionless response (MVDR-IC). Second, we propose to preserve the IC of the noise component using the binaural MWF with partial noise estimation (MWF-N) and the binaural MVDR beamformer with partial noise estimation (MVDR-N), for which closed-form expressions exist. In addition, we show that for the MVDR-N a closed-form expression can be derived for the tradeoff parameter yielding a desired magnitude squared coherence (MSC) for the output noise component. Since contrary to the MWF-IC and the MWF-N the MVDR-IC and the MVDR-N do not take into account the spectro-temporal properties of the speech and the noise components, we propose to apply a spectral postfilter to the filter outputs, improving the noise reduction performance. The performance of all algorithms is compared in several diffuse noise scenarios. The simulation results show that both the MVDR-IC and the MVDR-N are able to preserve the MSC of the noise component, while generally the MVDR-IC shows a slightly better noise reduction performance at a larger complexity. Further, simulation results show that applying a spectral postfilter leads to a very similar performance for all considered algorithms in terms of noise reduction and speech distortion. Daniel Marquardt, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Two-Microphone Hearing Aids Using Prediction Error Method for Adaptive Feedback ControlabstractA challenge in hearing aids is adaptive feedback control which often uses an adaptive filter to estimate the feedback path. This estimate of the feedback path usually results in a bias due to the correlation between the loudspeaker signal and the incoming signal. The prediction error method (PEM) is a popular method for reducing this bias for adaptive feedback control (AFC) in hearing aids, providing a significant performance improvement compared to conventional adaptive feedback control techniques. However, the PEM-based AFC (PEM-AFC) applications are still limited to single-microphone single-loudspeaker (SMSL) systems. This paper investigates the application of the PEM-AFC to a two-microphone single-loudspeaker hearing aid with detailed theoretical analysis as well as practical experiments. In the proposed method, PEM-AFC2, we use the two-microphone adaptive feedback control (AFC2) method with two microphones and one loudspeaker. The incoming signals at the two microphones are related by a relative transfer function (RTF) which is used to predict the incoming signal at the main microphone. In addition, a prefilter is employed to prewhiten the loudspeaker and the microphone signals before the adaptive filter estimates. As a result, the proposed method obtains a lower bias and a faster tracking rate compared to the PEM-AFC and the AFC2 method, while still maintaining a good quality of the incoming signal. A new derivation for optimal filters in the AFC2 method will also be provided. The performance of the proposed method is evaluated for speech shaped noise as incoming signal and with undermodeling the RTF as well as with perfect modeling the RTF. Moreover, different types of incoming signals and a sudden change of feedback paths are also considered. The experimental results show that the proposed approach yields a significant performance improvement compared to existing state-of-the-art AFC methods such as the PEM-AFC and the AFC2. Linh Thi Thuc Tran, Sven Nordholm, Henning F. Schepker, Hai Huyen Dam, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Comparison of two binaural beamforming approaches for hearing aidsabstractBeamforming algorithms in binaural hearing aids are crucial to improve speech understanding in background noise for hearing impaired persons. In this study, we compare and evaluate the performance of two recently proposed minimum variance (MV) beamforming approaches for binaural hearing aids. The binaural linearly constrained MV (BLCMV) beamformer applies linear constraints to maintain the target source and mitigate the interfering sources, taking into account the reverberant nature of sound propagation. The inequality constrained MV (ICMV) beamformer applies inequality constraints to maintain the target source and mitigate the interfering sources, utilizing estimates of the direction of arrivals (DOAs) of the target and interfering sources. The similarities and differences between these two approaches is discussed and the performance of both algorithms is evaluated using simulated data and using real-world recordings, particularly focusing on the robustness to estimation errors of the relative transfer functions (RTFs) and DOAs. The BLCMV achieves a good performance if the RTFs are accurately estimated while the ICMV shows a good robustness to DOA estimation errors. Elior Hadad, Daniel Marquardt, Wenqiang Pu, Sharon Gannot, Simon Doclo, Zhi-Quan Luo, Ivo Merks, Tao Zhang 0024 |
ICASSP | 5 |
| 2017 | Measuring, modelling and predicting perceived reverberationabstractThis paper investigates the relationship between the perceived level of reverberation and parameters measured from the room impulse response (RIR), as well as the design of an instrumental measure that predicts this perceived level. We first present the results of an experimental listening test conducted to assess the level of perceived reverberation in speech captured by a single microphone, before analysing the gathered data to assess the influence of parameters such as the reverberation time (T60) or the direct-to-reverberant ratio (DRR). Secondly, we use the results of this analysis to improve the signal based reverberation decay tail (RDT) measure, previously proposed by the authors to predict the perceived level of reverberation. The accuracy of the proposed measure is evaluated in terms of correlation with the subjective scores and compared to the performance of predictors using parameters extracted from the RIR. Results show that the proposed modifications to the RDT does improve its accuracy. Though still slightly outperformed by measures based on parameters of the RIR, we believe the proposed measure to be useful in scenarios in which the RIR or its parameters are unknown. Hamza A. Javed, Benjamin Cauchi, Simon Doclo, Patrick A. Naylor, Stefan Goetze |
ICASSP | 3 |
| 2017 | Late reverberant power spectral density estimation based on an eigenvalue decompositionabstractMulti-channel methods for estimating the late reverberant power spectral density (PSD) rely on an estimate of the direction of arrival (DOA) of the speech source or of the relative early transfer functions (RETFs) of the target signal from a reference microphone to all microphones. The DOA and the RETFs may be difficult to estimate accurately, particularly in highly reverberant and noisy scenarios. In this paper we propose a novel multi-channel method to estimate the late reverberant PSD which does not require estimates of the DOA or RETFs. The late reverberation is modeled as an isotropic sound field and the late reverberant PSD is estimated based on the eigenvalues of the prewhitened received signal PSD matrix. Experimental results demonstrate the advantages of using the proposed estimator in a multi-channel Wiener filter for speech dereverberation, outperforming a recently proposed maximum likelihood estimator both when the DOA is perfectly estimated as well as in the presence of DOA estimation errors. Ina Kodrasi, Simon Doclo |
ICASSP | 2 |
| 2017 | Null-steering beamformer for acoustic feedback cancellation in a multi-microphone earpiece optimizing the maximum stable gainabstractCommonly adaptive filters are used to reduce the acoustic feedback in hearing aids. While theoretically allowing for perfect cancellation of the feedback signal, in practice the adaptive filter solution is typically biased due to the closed-loop hearing aid system. In contrast to conventional behind-the-ear hearing aids, in this paper we consider an earpiece with multiple integrated microphones. For such an earpiece it has previously been proposed to use a fixed beamformer to reduce the acoustic feedback in the microphones which has been designed to minimize a least-squares cost function. In this paper we propose to design the beamformer by minimizing a min-max cost function which directly maximizes the maximum stable gain of the earpiece. Furthermore, we propose a robust extension of the min-max cost function maximizing the worst-case maximum stable gain over a set of acoustic feedback paths. Experimental results using measured acoustic feedback paths show that the feedback cancellation performance of the fixed beamformer can be considerably improved by minimizing the proposed min-max optimization problem, while maintaining a high perceptual quality of the incoming signal. Henning F. Schepker, Linh Thi Thuc Tran, Sven Nordholm, Simon Doclo |
ICASSP | 4 |
| 2017 | Proportionate NLMS for adaptive feedback control in hearing aidsabstractThe proportionate normalized least-mean-squares (PNLMS) algorithm is commonly used in acoustic echo cancellation (AEC) context. It provides faster initial convergence and tracking rates compared to the NLMS algorithm for the case of sparse echo impulse responses. The improved PNLMS algorithm (IPNLMS) has been proven to be more powerful than PNLMS by exploiting new rules for computing the weight of each step-size corresponding to each adaptive filter coefficient. However, the application of the PNLMS and the IPNLMS algorithms for adaptive feedback control (AFC) in hearing aids (HAs) is still limited due to high correlation between the loudspeaker and incoming signals. This paper proposes implementations of the PNLMS/IPNLMS algorithms for AFC using the prediction error method (PEM) for hearing aids. The proposed methods have been evaluated for both speech and music incoming signals. Simulation shows that the proposed methods have faster initial convergence and tracking than the PEM using the NLMS algorithm (PEM-NLMS). Linh Thi Thuc Tran, Henning F. Schepker, Simon Doclo, Hai Huyen Dam, Sven Nordholm |
ICASSP | 3 |
| 2017 | EEG-based auditory attention decoding: Impact of reverberation, noise and interference reductionabstractTo identify the attended speaker from single-trial EEG recordings in an acoustic scenario with two competing speakers, an auditory attention decoding (AAD) method has recently been proposed. The AAD method requires the clean speech signals of both the attended and the unattended speaker as reference signals for decoding. However, in practice only the binaural signals, containing several undesired acoustic components (reverberation, background noise and interference), and influenced by anechoic head-related transfer functions (HRTFs), are available. To generate appropriate reference signals for decoding from the binaural signals, it is important to understand the impact of these acoustic components on the AAD performance. In this paper, we investigate this impact for decoding several acoustic conditions (anechoic, reverberant, noisy, and reverberant-noisy) by using simulated speech signals in which different acoustic components have been reduced. The experimental results show that for obtaining a good decoding performance the joint suppression of reverberation, background noise and interference as undesired acoustic components is of great importance. Ali Aroudi, Simon Doclo |
SMC | 2 |
| 2017 | Adaptive Speech Dereverberation Using Constrained Sparse Multichannel Linear PredictionabstractIn this letter, we present an adaptive speech dereverberation method based on constrained sparse multichannel linear prediction (MCLP), minimizing the mixed ℓ2,pnorm of the desired component. In order to prevent overestimation of the undesired reverberant component, possibly leading to severe distortions of the output, we propose to use a statistical model for late reverberation to limit the power of the MCLP-based estimate. The resulting constrained optimization problem is solved by using the alternating direction method of multipliers, resulting in two variants of the dereverberation algorithm. Simulation results show that the proposed constraint increases the robustness with respect to parameter selection and improves the usability for dynamic scenarios in comparison to the unconstrained method. Ante Jukic, Toon van Waterschoot, Simon Doclo |
IEEE Signal Process. Lett. | 3 |
| 2017 | Signal-Dependent Penalty Functions for Robust Acoustic Multi-Channel EqualizationabstractAcoustic multi-channel equalization techniques, which aim to achieve dereverberation by reshaping the room impulse responses (RIRs) between the source and the microphone array, are known to be highly sensitive to RIR perturbations. In order to increase the robustness against RIR perturbations, several signal-independent methods have been proposed, which only rely on the available perturbed RIRs and do not incorporate any knowledge about the output signal. This paper presents a novel signal-dependent method to increase the robustness of equalization techniques by enforcing the output signal to exhibit spectrotemporal characteristics of a clean speech signal. Motivated by the sparse nature of clean speech, we propose to extend the cost function of state-of-the-art least squares equalization techniques, i.e., the multiple-input/output inverse theorem (MINT), relaxed multi-channel least squares (RMCLS), and partial multi-channel equalization based on MINT (PMINT), with a signal-dependent penalty function promoting sparsity of the output signal in the short-time Fourier transform domain. Three conventionally used sparsity-promoting penalty functions are investigated, i.e., the l0-norm, the l1-norm, and the weighted l1-norm, and the sparsitypromoting reshaping filters are iteratively computed using the alternating direction method of multipliers. Simulation results for several acoustic systems and RIR perturbations demonstrate that incorporating sparsity-promoting penalty functions significantly increases the robustness of MINT, RMCLS, and PMINT, with the weighted l1-norm typically outperforming the l0-norm and the l1-norm. Furthermore, it is shown that the weighted l1-norm sparsity-promoting PMINT technique outperforms the other sparsity-promoting techniques in terms of perceptual speech quality. Finally, it is shown that the signal-dependent weighted l1-norm sparsity-promoting PMINT technique yields a similar or better dereverberation performance than the signal-independent regularized PMINT technique, confirming the advantage of using signal-dependent penalty functions for robust dereverberation filter design. Ina Kodrasi, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Correction to "Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise"abstractPresents corrections to the paper, "A Novel Approach Based on Marine Radar Data Analysis for High-Resolution Bathymetry Map Generation." Adam Kuklasinski, Simon Doclo, Søren Holdt Jensen, Jesper Rindom Jensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Auditory attention decoding with EEG recordings using noisy acoustic reference signalsabstractTo decode auditory attention from electroencephalography (EEG) recordings in a cocktail-party scenario with two competing speakers a least-squares method has recently been proposed, showing a promising decoding accuracy. This method however requires the clean speech signals of both the attended and the unattended speaker to be available as reference signals, which is difficult to achieve from the noisy recorded microphone signals in practice. In addition, optimizing the parameters involved in the spatio-temporal filter design is of crucial importance in order to reach the largest possible decoding performance. In this paper, the influence of noisy acoustic reference signals and the spatio-temporal filter and regularization parameters on the decoding performance is investigated. The results show that to some extent the decoding performance is robust to noisy acoustic reference signals, depending on the noise type. Furthermore, we demonstrate the crucial influence of several parameters on the decoding performance, especially when the acoustic reference signals used for decoding have been corrupted by noise. Ali Aroudi, Bojana Mirkovic, Maarten De Vos, Simon Doclo |
ICASSP | 4 |
| 2016 | Perceptual and instrumental evaluation of the perceived level of reverberationabstractPerceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significant. Therefore, an efficient perceptual measure of the perceived level of reverberation is needed. We compare the use of a multiple stimuli test with the use of pairwise comparison for the evaluation of the perceived level of reverberation. The results suggest that using multiple stimuli is preferable to pairwise comparison as long as the number of conditions to be compared is not too large. Additionally, we use the results from the conducted perceptual measurements to examine the reliability of existing instrumental measures of the perceived level of reverberation. Our observations show which instrumental measures are effective in highlighting differences between RIR characteristics and which ones have to be preferred if one aims at predicting the level of reverberation perceived by a human assessor. Benjamin Cauchi, Hamza A. Javed, Timo Gerkmann, Simon Doclo, Stefan Goetze, Patrick A. Naylor |
ICASSP | 4 |
| 2016 | Extensions of the binaural MWF with interference reduction preserving the binaural cues of the interfering sourceabstractRecently, an extension of the binaural multichannel Wiener filter (BMWF), referred to as BMWF-IRo, was presented in which an interference rejection constraint was added to the BMWF cost function. Although the BMWF-IRo aims to entirely suppress the interfering source, residual interfering sources (as well as unconstrained noise sources) are undesirably perceived as impinging the array from the desired source direction. In this paper, we propose two extensions of the BMWF-IRo that address this issue by preserving the spatial impression of the interfering source. In the first extension, the binaural cues of the interfering source are preserved, while those of the desired source may be slightly distorted. In the second extension, the binaural cues of both the desired and interfering sources are preserved. Simulation results show that the noise reduction performance of both proposed extensions is comparable to the BMWF-, IRo. Elior Hadad, Daniel Marquardt, Simon Doclo, Sharon Gannot |
ICASSP | 3 |
| 2016 | Robust sparsity-promoting acoustic multi-channel equalization for speech dereverberationabstractThis paper presents a novel signal-dependent method to increase the robustness of acoustic multi-channel equalization techniques against room impulse response (RIR) estimation errors. Aiming at obtaining an output signal which better resembles a clean speech signal, we propose to extend the acoustic multi-channel equalization cost function with a penalty function which promotes sparsity of the output signal in the short-time Fourier transform domain. Two conventionally used sparsity-promoting penalty functions are investigated, i.e., the l0-norm and the l1-norm, and the sparsity-promoting filters are iteratively computed using the alternating direction method of multipliers. Simulation results for several RIR estimation errors show that incorporating a sparsity-promoting penalty function significantly increases the robustness, with the l1-norm penalty function outperforming the l0-norm penalty function. Ina Kodrasi, Ante Jukic, Simon Doclo |
ICASSP | 3 |
| 2016 | Maximum likelihood PSD estimation for speech enhancement in reverberant and noisy conditionsabstractWe propose a novel Power Spectral Density (PSD) estimator for multi-microphone systems operating in reverberant and noisy conditions. The estimator is derived using the maximum likelihood approach and is based on a blocked and pre-whitened additive signal model. The intended application of the estimator is in speech enhancement algorithms, such as the Multi-channel Wiener Filter (MWF) and the Minimum Variance Distortionless Response (MVDR) beamformer. We evaluate these two algorithms in a speech dereverberation task and compare the performance obtained using the proposed and a competing PSD estimator. Instrumental performance measures indicate an advantage of the proposed estimator over the competing one. In a speech intelligibility test all algorithms significantly improved the word intelligibility score. While the results suggest a minor advantage of using the proposed PSD estimator, the difference between algorithms was found to be statistically significant only in some of the experimental conditions. Adam Kuklasinski, Simon Doclo, Jesper Jensen 0001 |
ICASSP | 2 |
| 2016 | Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aidsabstractBesides noise reduction, an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. For the desired speech source and an interfering source, e.g., competing speaker, this can be achieved by preserving their relative transfer functions (RTFs). It has been shown that the binaural multi-channel Wiener filter (MWF) preserves the RTF of the desired speech source, but typically distorts the RTF of the interfering source. To this end, in this paper we propose an extension of the binaural MWF, i.e. the binaural MWF with RTF preservation (MWF-RTF) aiming to preserve the RTF of the interfering source. Analytical expressions for the performance of the binaural MWF and the MWF-RTF in terms of noise reduction and binaural cue preservation are derived, using which their performance is thoroughly compared. Simulation results using binaural behind-the-ear impulse responses measured in a reverberant environment validate the derived analytical expressions, showing that the MWF-RTF yields a better performance than the binaural MWF in terms of the signal-to-interference ratio and binaural cue preservation of the interfering source, while the overall noise reduction performance is slightly degraded. Daniel Marquardt, Elior Hadad, Sharon Gannot, Simon Doclo |
ICASSP | 4 |
| 2016 | Improving adaptive feedback cancellation in hearing aids using an affine combination of filtersabstractIn adaptive feedback cancellation an adaptive filter is used to model the acoustic feedback path between the hearing aid loudspeaker and the microphone. An important parameter for adaptive filters is the step-size, providing a trade-off between fast convergence and low steady-state misalignment. In order to achieve both fast convergence as well as low steady-state misalignment, it has been proposed to use an affine combination scheme of two filters operating with different step-sizes. In this paper we apply such an affine combination scheme to the acoustic feedback cancellation problem in hearing aids. We show that for speech signals a time-domain affine combination scheme yields a biased solution. To reduce this bias we propose to use a partitioned-block frequency-domain affine combination scheme. Experimental results using measured acoustic feedback paths show that in terms of misalignment and added stable gain the proposed adaptive feedback cancellation system outperforms a system that only uses a single adaptive filter with either of the fixed step-sizes used for the affine combination scheme. Henning F. Schepker, Linh Thi Thuc Tran, Sven Nordholm, Simon Doclo |
ICASSP | 4 |
| 2016 | The Binaural LCMV Beamformer and its Performance AnalysisabstractThe recently proposed binaural linearly constrained minimum variance (BLCMV) beamformer is an extension of the well-known binaural minimum variance distortionless response (MVDR) beamformer, imposing constraints for both the desired and the interfering sources. Besides its capabilities to reduce interference and noise, it also enables to preserve the binaural cues of both the desired and interfering sources, hence making it particularly suitable for binaural hearing aid applications. In this paper, a theoretical analysis of the BLCMV beamformer is presented. In order to gain insights into the performance of the BLCMV beamformer, several decompositions are introduced that reveal its capabilities in terms of interference and noise reduction, while controlling the binaural cues of the desired and the interfering sources. When setting the parameters of the BLCMV beamformer, various considerations need to be taken into account, e.g. based on the amount of interference and noise reduction and the presence of estimation errors of the required relative transfer functions (RTFs). Analytical expressions for the performance of the BLCMV beamformer in terms of noise reduction, interference reduction, and cue preservation are derived. Comprehensive simulation experiments, using measured acoustic transfer functions as well as real recordings on binaural hearing aids, demonstrate the capabilities of the BLCMV beamformer in various noise environments. Elior Hadad, Simon Doclo, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Joint Dereverberation and Noise Reduction Based on Acoustic Multi-Channel EqualizationabstractRegularized acoustic multi-channel equalization techniques, such as regularized partial multi-channel equalization based on the multiple-input/output inverse theorem (RPMINT), are able to achieve a high dereverberation performance in the presence of room impulse response perturbations but may lead to amplification of the additive noise. In this paper, two time-domain techniques aiming at joint dereverberation and noise reduction based on acoustic multi-channel equalization are proposed. The first technique, namely RPMINT for joint dereverberation and noise reduction (RPM-DNR), extends RPMINT by explicitly taking the noise statistics into account. In addition to the regularization parameter used in RPMINT, the RPM-DNR technique introduces an additional weighting parameter, enabling a trade-off between dereverberation and noise reduction. The second technique, namely multi-channel Wiener filter for joint dereverberation and noise reduction (MWF-DNR), takes both the speech and the noise statistics into account and uses the RPMINT filter to compute a dereverberated reference signal for the multi-channel Wiener filter. The MWF-DNR technique also introduces an additional weighting parameter, which now provides a trade-off between speech distortion and noise reduction. To automatically select the regularization and weighting parameters, for the RPM-DNR technique a novel procedure based on the L-hypersurface is proposed, whereas for the MWF-DNR technique two decoupled optimization procedures based on the L-curve are used. Extensive simulations demonstrate using instrumental measures that the RPM-DNR technique maintains the dereverberation performance of the RPMINT technique while improving its noise reduction performance. Furthermore, it is shown that the MWF-DNR technique yields a significantly better noise reduction performance than the RPM-DNR technique at the expense of a worse dereverberation performance. Ina Kodrasi, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and NoiseabstractIn this contribution, we focus on the problem of power spectral density (PSD) estimation from multiple microphone signals in reverberant and noisy environments. The PSD estimation method proposed in this paper is based on the maximum likelihood (ML) methodology. In particular, we derive a novel ML PSD estimation scheme that is suitable for sound scenes which besides speech and reverberation consists of an additional noise component whose second-order statistics are known. The proposed algorithm is shown to outperform an existing similar algorithm in terms of PSD estimation accuracy. Moreover, it is shown numerically that the mean-squared estimation error achieved by the proposed method is near the limit set by the corresponding Cramér-Rao lower bound. The speech dereverberation performance of a multichannel Wiener filter based on the proposed PSD estimators is measured using several instrumental measures and is shown to be higher than when the competing estimator is used. Moreover, we perform a speech intelligibility test where we demonstrate that both the proposed and the competing PSD estimators lead to similar intelligibility improvements. Adam Kuklasinski, Simon Doclo, Søren Holdt Jensen, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Speech Dereverberation Using Non-Negative Convolutive Transfer Function and Spectro-Temporal ModelingabstractThis paper presents two single-channel speech dereverberation methods to enhance the quality of speech signals that have been recorded in an enclosed space. For both methods, the room acoustics are modeled using a non-negative approximation of the convolutive transfer function (N-CTF), and to additionally exploit the spectral properties of the speech signal, such as the low-rank nature of the speech spectrogram, the speech spectrogram is modeled using non-negative matrix factorization (NMF). Two methods are described to combine the N-CTF and NMF models. In the first method, referred to as the integrated method, a cost function is constructed by directly integrating the speech NMF model into the N-CTF model, while in the second method, referred to as the weighted method, the N-CTF and NMF based cost functions are weighted and summed. Efficient update rules are derived to solve both optimization problems. In addition, an extension of the integrated method is presented, which exploits the temporal dependencies of the speech signal. Several experiments are performed on reverberant speech signals with and without background noise, where the integrated method yields a considerably higher speech quality than the baseline N-CTF method and a state-of-the-art spectral enhancement method. Moreover, the experimental results indicate that the weighted method can even lead to a better performance in terms of instrumental quality measures, but that the optimal weighting parameter depends on the room acoustics and the utilized NMF model. Modeling the temporal dependencies in the integrated method was found to be useful only for highly reverberant conditions. Nasser Mohammadiha, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Regularization Approaches for Synthesizing HRTF Directivity PatternsabstractAs an alternative to traditional artificial heads, it is possible to synthesize individual head-related transfer functions (HRTFs) using a so-called virtual artificial head (VAH), consisting of a microphone array with an appropriate topology and filter coefficients optimized using a narrowband least squares cost function. The resulting spatial directivity pattern of such a VAH is known to be sensitive to small deviations of the assumed microphone characteristics, e.g., gain, phase and/or the positions of the microphones. In many beamformer design procedures, this sensitivity is reduced by imposing a white noise gain (WNG) constraint on the filter coefficients for a single desired look direction. In this paper, this constraint is shown to be inappropriate for regularizing the HRTF synthesis with multiple desired directions and three alternative different regularization approaches are proposed and evaluated. In the first approach, the measured deviations of the microphone characteristics are taken into account in the filter design. In the second approach, the filter coefficients are regularized using the mean WNG for all directions. The third approach additionally takes into account several frequency bins into both the optimization and the regularization. The different proposed regularization approaches are compared using analytic and measured transfer functions, including random deviations. Experimental results show that the approach using multiple frequency bands mimicking the spectral resolution of the human auditory system yields the best robustness among the considered regularization approaches. Eugen Rasumow, Martin Hansen, Steven van de Par, Dirk Puschel, Volker Mellert, Simon Doclo, Matthias Blau |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | A Semidefinite Programming Approach to Min-max Estimation of the Common Part of Acoustic Feedback Paths in Hearing AidsabstractThe convergence speed and the computational complexity of adaptive feedback cancellation algorithms both depend on the number of adaptive parameters used to model the acoustic feedback paths. To reduce the number of adaptive parameters it has been proposed to decompose the acoustic feedback paths as the convolution of a time-invariant common part and time-varying variable parts. Instead of estimating all parameters of the common and variable parts by minimizing the misalignment using a least-squares cost function, in this paper we propose to formulate the parameter estimation problem as a min-max optimization problem aiming to maximize the maximum stable gain (MSG). We formulate the min-max optimization problem as a semidefinite program and use a constraint based on Lyapunov theory to guarantee stability of the estimated common pole-zero filter. Experimental results using measured acoustic feedback paths show that the proposed min-max optimization outperforms least-squares optimization in terms of the MSG. Furthermore, the results indicate that the proposed common part decomposition is able to increase the MSG and reduce the number of variable part parameters even for unknown feedback paths that were not included in the optimization. Simulation results using an adaptive feedback cancellation algorithm based on the prediction-error-method show that the convergence speed can be increased by using the proposed feedback path decomposition. Henning F. Schepker, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Least-Squares Estimation of the Common Pole-Zero Filter of Acoustic Feedback Paths in Hearing AidsabstractIn adaptive feedback cancellation both the convergence speed and the computational complexity depend on the number of adaptive parameters used to model the acoustic feedback paths. To reduce the number of adaptive parameters, it has been proposed to model the acoustic feedback paths as the convolution of a time-invariant common pole-zero filter and time-varying all-zero filters, enabling to track fast changes. In this paper, a novel procedure to estimate the common pole-zero filter of acoustic feedback paths is presented. In contrast to previous approaches which minimize the so-called equation-error, we propose to approximate the desired output-error minimization by employing a weighted least-squares procedure motivated by the Steiglitz-McBride iteration. The estimation of the common pole-zero filter is formulated as a semidefinite programming problem, to which a constraint based on the Lyapunov theory is added in order to guarantee the stability of the estimated pole-zero filter. Experimental results using measured acoustic feedback paths from a two microphone behind-the-ear hearing aid show that the proposed optimization procedure using the Lyapunov constraint outperforms existing optimization procedures in terms of modelling accuracy and added stable gain. Henning F. Schepker, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Correlation Maximization-Based Sampling Rate Offset Estimation for Distributed Microphone ArraysabstractIn this paper, we investigate the sampling rate mismatch problem in distributed microphone arrays and propose a correlation maximization algorithm to blindly estimate the sampling rate offset between two asynchronously sampled microphone signals. We approximate the sampling rate offset with a linear-phase drift model in the short-time Fourier transform (STFT) domain and show that the correlation coefficient between two microphone signals tends to present the highest value when the sampling of the two microphone signals is synchronized. Based on this finding we propose the correlation maximization algorithm, which performs sampling rate compensation on two microphone signals with different possible offset values and calculates their correlation coefficient after compensation. The offset value that leads to the largest correlation coefficient is chosen as the optimal estimate. Since the precision of the STFT linear-phase drift model used in the algorithm degrades as the sampling rate offset or the signal length is increased, we further propose a two-stage exhaustive search scheme to detect the optimal sampling rate offset. This scheme is able to minimize the influence of the linear-phase drift model error in order to improve the sampling rate offset estimation accuracy. Both simulated as well as real-world experiments confirm the effectiveness of the proposed algorithm. Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | On application of non-negative matrix factorization for ad hoc microphone array calibration from incomplete noisy distancesabstractWe propose to use non-negative matrix factorization (NMF) to estimate the unknown pairwise distances and reconstruct a distance matrix for microphone array position calibration. We develop new multiplicative update rules for NMF with incomplete input matrix that take into account the symmetry of the distance matrix. Additionally, we develop a convex matrix completion method which is related to an l2-regularized symmetric NMF. Thorough experiments demonstrate that the proposed methods lead to substantial improvement over the state-of-the-art techniques in a wide range of signal-to-noise and unknown-distance ratios. The convex symmetric matrix completion method was found to be the most robust method with less computational cost. Afsaneh Asaei, Nasser Mohammadiha, Mohammad Javad Taghizadeh, Simon Doclo, Hervé Bourlard |
ICASSP | 4 |
| 2015 | Binaural multichannel Wiener filter with directional interference rejectionabstractIn this paper we consider an acoustic scenario with a desired source and a directional interference picked up by hearing devices in a noisy and reverberant environment. We present an extension of the binaural multichannel Wiener filter (BMWF), by adding an interference rejection constraint to its cost function, in order to combine the advantages of spatial and spectral filtering while mitigating directional interferences. We prove that this algorithm can be decomposed into the binaural linearly constrained minimum variance (BLCMV) algorithm followed by a single channel Wiener post-filter. The proposed algorithm yields improved interference rejection capabilities, as compared with the BMWF. Moreover, by utilizing the spectral information on the sources, it is demonstrating better SNR measures, as compared with the BLCMV. Elior Hadad, Daniel Marquardt, Simon Doclo, Sharon Gannot |
ICASSP | 3 |
| 2015 | Speaker change detection and speaker diarization using spatial informationabstractIn this paper, we present a novel speaker change detection and speaker diarization algorithm using spatial information in the form of features derived from estimated Room Impulse Response (RIR)s. A blind system identification approach is used to obtain an estimate of the RIRs, from which the C5 feature is derived and used in the labeling algorithm. Experimental results using 2 speakers for different locations within a fixed room show that our approach achieves a higher hit rate in the speaker change detection task and a lower variance in the diarization error rate when compared with a baseline algorithm. Mathieu Hu, Dushyant Sharma, Simon Doclo, Mike Brookes, Patrick A. Naylor |
ICASSP | 3 |
| 2015 | Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximationabstractIn many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transform domain, which is typically performed in each frequency bin independently without taking into account the spectral structure of the speech signal. Since it is widely accepted that a speech spectrogram can be well approximated with a low-rank matrix, e.g., using a spectral dictionary, in this paper we propose to incorporate a low-rank matrix approximation of the speech spectrogram into the MCLP-based speech dereverberation. The low-rank approximation is obtained using nonnegative matrix factorization with Itakura-Saito divergence. Experimental results for several measured acoustic systems show that incorporating a low-rank approximation improves the dereverberation performance in terms of instrumental speech quality measures. Ante Jukic, Nasser Mohammadiha, Toon van Waterschoot, Timo Gerkmann, Simon Doclo |
ICASSP | 5 |
| 2015 | Curvature-based optimization of the trade-off parameter in the speech distortion weighted multichannel wiener filterabstractThe objective of the speech distortion weighted multichannel Wiener filter (MWF) is to reduce background noise while controlling speech distortion. This can be achieved by means of a trade-off parameter, hence, selecting an optimal trade-off parameter is of crucial importance. Aiming at incorporating knowledge about the resulting speech distortion and noise power, in this paper we propose to compute the trade-off parameter as the point of maximum curvature of the parametric plot of noise power versus speech distortion. To determine a narrowband trade-off parameter, an analytical expression is derived for computing the point of maximum curvature, whereas to determine a broadband parameter an optimization routine is used. The speech distortion and the noise power terms can also be weighted in advance, e.g. based on perceptually motivated criteria. Experimental results show that using the proposed method instead of the MWF improves the intelligibility weighted SNR without significantly degrading the speech distortion. Ina Kodrasi, Daniel Marquardt, Simon Doclo |
ICASSP | 3 |
| 2015 | Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparisonabstractIn this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities (PSDs). We first derive closedform expressions for the mean square error (MSE) of both PSD estimators and then show that one estimator - previously used for speech dereverberation by the authors - always yields a better MSE. Only in the case of a two microphone array or for special spatial distributions of the interference both estimators yield the same MSE. The theoretically derived MSE values are in good agreement with numerical simulation results and with instrumental speech quality measures in a realistic speech dereverberation task for binaural hearing aids. Adam Kuklasinski, Simon Doclo, Timo Gerkmann, Søren Holdt Jensen, Jesper Jensen 0001 |
ICASSP | 2 |
| 2015 | Interaural coherence preservation in MWF-based binaural noise reduction algorithms using partial noise estimationabstractBesides noise reduction an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of both desired and undesired sound sources. Recently an extension of the binaural Multi-channel Wiener filter (MWF), namely the MWF-IC, has been presented which aims to preserve the Interaural Coherence (IC) of the noise component. Since for the MWF-IC no closed-form solution exists, in this paper we propose to preserve the IC using the binaural MWF with partial noise estimation (MWF-N), for which a closed-form solution exists. Furthermore, we derive a closed-form expression for the trade-off parameter in the MWF-N yielding a predefined IC at the filter output. Experimental results in a diffuse noise scenario show that both the MWF-IC and the MWF-N preserve the IC of the output noise component. However, the MWF-IC yields a better noise reduction performance whereas the MWF-N introduces less speech distortion. Daniel Marquardt, Volker Hohmann, Simon Doclo |
ICASSP | 3 |
| 2015 | Joint acoustic and spectral modeling for speech dereverberation using non-negative representationsabstractThis paper proposes a single-channel speech dereverberation method enhancing the spectrum of the reverberant speech signal. The proposed method uses a non-negative approximation of the convolutive transfer function (N-CTF) to simultaneously estimate the magnitude spectrograms of the speech signal and the room impulse response (RIR). To utilize the speech spectral structure, we propose to model the speech spectrum using non-negative matrix factorization, which is directly used in the N-CTF model resulting in a new cost function. We derive new estimators for the parameters by minimizing the obtained cost function. Additionally, to investigate the effect of the speech temporal dynamics for dereverberation, we use a frame stacking method and derive optimal estimators. Experiments are performed for two measured RIRs and the performance of the proposed method is compared to the performance of a state-of-the-art dereverberation method enhancing the speech spectrum. Experimental results show that the proposed method improved instrumental speech quality measures, where using speech temporal dynamics was found to be beneficial in severe reverberation conditions. Nasser Mohammadiha, Paris Smaragdis, Simon Doclo |
ICASSP | 3 |
| 2015 | Individualizing a monaural beamformer for cochlear implant usersabstractSpeech intelligibility in noisy environments is still quite limited for cochlear implant (CI) users. Classical beamformers such as the Generalized Sidelobe Canceller (GSC) can provide large improvements in speech intelligibility for CI users. These algorithms have been adopted from hearing aids and multimedia applications into the CI field. However, their optimization taking into consideration the peculiarities of electrical hearing with a CI has not yet been completely investigated. This paper presents a novel method to optimize the performance of a GSC for each individual CI user. We show through a combination of objective and novel subjective measures, how much distortion can be tolerated by a CI user without decreasing speech intelligibility. Experimental results with 5 CI users show that a GSC delivering just noticeable distortion is the one maximizing speech intelligibility for CI users. Waldo Nogueira, Marta Lopez 0003, Thilo Rode, Simon Doclo, Andreas Büchner |
ICASSP | 4 |
| 2015 | Common part estimation of acoustic feedback paths in hearing aids optimizing maximum stable gainabstractThe computational complexity and convergence speed of adaptive feedback cancellation algorithms depend on the number of adaptive parameters used to model the acoustic feedback path. To reduce the number of adaptive parameters it has been proposed to decompose the acoustic feedback path as the convolution of a (time-invariant) common part and a (time-varying) variable part. Typically the problem of estimating all the required coefficients has been formulated as a least-squares optimization problem. In contrast, in this paper we propose to formulate the estimation problem as a minmax optimization problem and show how this is associated with the maximum stable gain of a hearing aid. Experimental results using measured acoustic feedback paths from a two-microphone behind-the-ear hearing aid show that the proposed minmax optimization outperforms the least-squares optimization in terms of maximum stable gain. Furthermore, the robustness of proposed common part decomposition for different feedback paths is evaluated. Henning F. Schepker, Simon Doclo |
ICASSP | 2 |
| 2015 | Model-based adaptive pre-processing of speech for enhanced intelligibility in noise and reverberation
Jan Rennies, Andreas Volgenandt, Henning F. Schepker, Simon Doclo |
INTERSPEECH | 4 |
| 2015 | Model-based integration of reverberation for noise-adaptive near-end listening enhancementabstractSpeech intelligibility is an important factor for successful speech communication in today’s society. So-called near-end listening enhancement (NELE) algorithms aim at improving speech intelligibility in conditions where the (clean) speech signal is accessible and can be modified prior to its presentation. However, many of these algorithms only consider the detrimental effect of noise and disregard the effect of reverberation. Therefore, in this paper we propose to additionally incorporate the detrimental effects of reverberation into noise-adaptive nearend listening enhancement algorithms. Based on the Speech Transmission Index (STI), which is widely used for speech intelligibility prediction, the effect of reverberation is effectively accounted for as an additional noise power term. This combined noise power term is used in a state-of-the-art noise-adaptive NELE algorithm. Simulations using two objective measures, the STI and the short-time objective intelligibility (STOI) measure demonstrate the potential of the proposed approach to improve the predicted speech intelligibility in noisy and reverberant conditions. Index Terms: speech-in-noise enhancement, reverberation, speech intelligibility, near-end listening enhancement Henning F. Schepker, David Hülsmeier, Jan Rennies, Simon Doclo |
INTERSPEECH | 4 |
| 2015 | Special issue on wireless acoustic sensor networks and ad hoc microphone arrays
Alexander Bertrand, Simon Doclo, Sharon Gannot, Nobutaka Ono, Toon van Waterschoot |
Signal Process. | 2 |
| 2015 | Analysis of the average performance of the multi-channel Wiener filter for distributed microphone arrays using statistical room acoustics
Toby Christian Lawin-Ore, Simon Doclo |
Signal Process. | 2 |
| 2015 | Theoretical Analysis of Binaural Transfer Function MVDR Beamformers with Interference Cue Preservation ConstraintsabstractThe objective of binaural noise reduction algorithms is not only to selectively extract the desired speaker and to suppress interfering sources (e.g., competing speakers) and ambient background noise, but also to preserve the auditory impression of the complete acoustic scene. For directional sources this can be achieved by preserving the relative transfer function (RTF) which is defined as the ratio of the acoustical transfer functions relating the source and the two ears and corresponds to the binaural cues. In this paper, we theoretically analyze the performance of three algorithms that are based on the binaural minimum variance distortionless response (BMVDR) beamformer, and hence, process the desired source without distortion. The BMVDR beamformer preserves the binaural cues of the desired source but distorts the binaural cues of the interfering source. By adding an interference reduction (IR) constraint, the recently proposed BMVDR-IR beamformer is able to preserve the binaural cues of both the desired source and the interfering source. We further propose a novel algorithm for preserving the binaural cues of both the desired source and the interfering source by adding a constraint preserving the RTF of the interfering source, which will be referred to as the BMVDR-RTF beamformer. We analytically evaluate the performance in terms of binaural signal-to-interference-and-noise ratio (SINR), signal-to-interference ratio (SIR), and signal-to-noise ratio (SNR) of the three considered beamformers. It can be shown that the BMVDR-RTF beamformer outperforms the BMVDR-IR beamformer in terms of SINR and outperforms the BMVDR beamformer in terms of SIR. Among all beamformers which are distortionless with respect to the desired source and preserve the binaural cues of the interfering source, the newly proposed BMVDR-RTF beamformer is optimal in terms of SINR. Simulations using acoustic transfer functions measured on a binaural hearing aid validate our theoretical results. Elior Hadad, Daniel Marquardt, Simon Doclo, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Multi-Channel Linear Prediction-Based Speech Dereverberation With Sparse PriorsabstractThe quality of speech signals recorded in an enclosure can be severely degraded by room reverberation. In this paper, we focus on a class of blind batch methods for speech dereverberation in a noiseless scenario with a single source, which are based on multi-channel linear prediction in the short-time Fourier transform domain. Dereverberation is performed by maximum-likelihood estimation of the model parameters that are subsequently used to recover the desired speech signal. Contrary to the conventional method, we propose to model the desired speech signal using a general sparse prior that can be represented in a convex form as a maximization over scaled complex Gaussian distributions. The proposed model can be interpreted as a generalization of the commonly used time-varying Gaussian model. Furthermore, we reformulate both the conventional and the proposed method as an optimization problem with an lp-norm cost function, emphasizing the role of sparsity in the considered speech dereverberation methods. Experimental evaluation in different acoustic scenarios show that the proposed approach results in an improved performance compared to the conventional approach in terms of instrumental measures for speech quality. Ante Jukic, Toon van Waterschoot, Timo Gerkmann, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Interaural Coherence Preservation in Multi-Channel Wiener Filtering-Based Noise Reduction for Binaural Hearing AidsabstractBesides noise reduction an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. To this end, an extension of the binaural multi-channel Wiener filter (MWF), namely the MWF-ITF, has been proposed, which aims to preserve the Interaural Transfer Function (ITF) of the noise sources. However, the MWF-ITF is well-suited only for directional noise sources but not for, e.g., spatially isotropic noise, whose spatial characteristics cannot be properly described by the ITF but rather by the Interaural Coherence (IC). Hence, another extension of the binaural MWF, namely the MWF-IC, has been recently proposed, which aims to preserve the IC of the noise component. Since for the MWF-IC a substantial tradeoff between noise reduction and IC preservation exists, in this paper we propose a perceptually constrained version of the MWF-IC, where the amount of IC preservation is controlled based on the IC discrimination ability of the human auditory system. In addition, a theoretical analysis of the binaural cue preservation capabilities of the binaural MWF and the MWF-ITF for spatially isotropic noise fields is provided. Several simulations in diffuse noise scenarios show that the perceptually constrained MWF-IC yields a controllable preservation of the IC without significantly degrading the output SNR compared to the binaural MWF and the MWF-ITF. Furthermore, contrary to the binaural MWF and MWF-ITF, the proposed algorithm retains the spatial separation between the output speech and noise components while the binaural cues of the speech component are only slightly distorted, such that the binaural hearing advantage for speech intelligibility can still be exploited. Daniel Marquardt, Volker Hohmann, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Theoretical Analysis of Linearly Constrained Multi-Channel Wiener Filtering Algorithms for Combined Noise Reduction and Binaural Cue Preservation in Binaural Hearing AidsabstractBesides noise reduction, an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. For the desired speech source and the interfering sources, e.g., competing speakers, this can be achieved by preserving their relative transfer functions (RTFs). It has been shown that the binaural multi-channel Wiener filter (MWF) preserves the RTF of the desired speech source, but typically distorts the RTF of the interfering sources. To this end, in this paper we propose two extensions of the binaural MWF, i.e., the binaural MWF with RTF preservation (MWF-RTF) aiming to preserve the RTF of the interfering source and the binaural MWF with interference rejection (MWF-IR) aiming to completely suppress the interfering source. Analytical expressions for the performance of the binaural MWF, MWF-RTF and MWF-IR in terms of noise reduction, speech distortion and binaural cue preservation are derived, showing that the proposed extensions yield a better performance in terms of the signal-to-interference ratio and preservation of the binaural cues of the directional interference, while the overall noise reduction performance is degraded compared to the binaural MWF. Simulation results using binaural behind-the-ear impulse responses measured in a reverberant environment validate the derived analytical expressions for the theoretically achievable performance of the binaural MWF, MWF-RTF, and MWF-IR, showing that the performance highly depends on the position of the interfering source and the number of microphones. Furthermore, the simulation results show that the MWF-RTF yields a very similar overall noise reduction performance as the binaural MWF, while preserving the binaural cues of both the speech and the interfering source. Daniel Marquardt, Elior Hadad, Sharon Gannot, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Noise Power Spectral Density Estimation Using MaxNSR Blocking MatrixabstractIn this paper, a multi-microphone noise reduction system based on the generalized sidelobe canceller (GSC) structure is investigated. The system consists of a fixed beamformer providing an enhanced speech reference, a blocking matrix providing a noise reference by suppressing the target speech, and a single-channel spectral post-filter. The spectral post-filter requires the power spectral density (PSD) of the residual noise in the speech reference, which can in principle be estimated from the PSD of the noise reference. However, due to speech leakage in the noise reference, the noise PSD is overestimated, leading to target speech distortion. To minimize the influence of the speech leakage, a maximum noise-to-speech ratio (MaxNSR) blocking matrix is proposed, which maximizes the ratio between the noise and the speech leakage in the noise reference. The proposed blocking matrix can be computed from the generalized eigenvalue decomposition of the correlation matrix of the microphone signals and the noise coherence matrix, which is assumed to be time-invariant. Experimental results in both stationary and nonstationary diffuse noise fields show that the proposed algorithm outperforms existing blocking matrices in terms of target speech blocking ability, noise estimation and noise reduction performance. Timo Gerkmann, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Speech dereverberation using weighted prediction error with Laplacian model of the desired signalabstractReverberation has a considerable impact on the quality and intelligibility of captured speech signals. In this paper we present an approach for blind multi-microphone speech dereverberation based on the weighted prediction error method, where the reverberant observations are modeled using multi-channel linear prediction in the short-time Fourier transform domain. Instead of using the commonly employed Gaussian distribution for the desired speech signal, the proposed approach uses a Laplacian distribution which is known to be more accurate in modeling speech signals. Maximum-likelihood estimation is used for estimating the model parameters, leading to a linear programming optimization problem. Experimental results, obtained using measured impulse responses, indicate that the proposed approach could be used to improve the dereverberation performance compared to the classical technique. Ante Jukic, Simon Doclo |
ICASSP | 2 |
| 2014 | Frequency-domain single-channel inverse filtering for speech dereverberation: Theory and practiceabstractThe objective of single-channel inverse filtering is to design an inverse filter that achieves dereverberation while being robust to an inaccurate room impulse response (RIR) measurement or estimate. Since a stable and causal inverse filter typically does not exist, approximate time-domain inverse filtering techniques such as singlechannel least-squares (SCLS) have been proposed. However, besides being computationally expensive and often infeasible, SCLS generally leads to distortions in the output signal in the presence of RIR inaccuracies. In this paper, a theoretical analysis is initially provided, showing that the direct inversion of the acoustic transfer function in the frequency-domain generally yields instability and acausality issues. In order to resolve these issues, a novel frequency-domain inverse filtering technique is proposed that incorporates regularization and uses a single-channel speech enhancement scheme. Experimental results demonstrate that the proposed technique yields a higher dereverberation performance and has a significantly lower computational complexity compared to the SCLS technique. Ina Kodrasi, Timo Gerkmann, Simon Doclo |
ICASSP | 3 |
| 2014 | Perceptually motivated coherence preservation in multi-channel wiener filtering based noise reduction for binaural hearing aidsabstractBesides noise reduction an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of both desired and undesired sound sources. Recently, an extension of the binaural Multi-channel Wiener filter (MWF), namely the MWF-IC, has been presented, which aims to preserve the Interaural Coherence (IC) of the noise component. Since for the MWF-IC a substantial trade-off between noise reduction and IC preservation exists, in this paper we propose a perceptually constrained version of the MWF-IC, where the amount of IC preservation is controlled based on psychoacoustic criterias of the IC discrimination ability of the human auditory system. In addition, we present a simplified version of the MWF-IC, resulting in a decrease of computational complexity. Experimental results show that the perceptually motivated MWF-IC and its simplified version yield a very similar performance and the loss in intelligibility weighted output SNR compared to the binaural MWF can be limited to 0.5 dB, whereas the spatial separation between the output speech and noise component is increased leading to better perceptual results. Daniel Marquardt, Volker Hohmann, Simon Doclo |
ICASSP | 3 |
| 2014 | Modeling the common part of acoustic feedback paths in hearing aids using a pole-zero modelabstractIn adaptive feedback cancellation the computational complexity and the convergence speed are determined by the number of adaptive parameters used to model the acoustic feedback path. Therefore it has been proposed to reduce the number of adaptive parameters by modeling the feedback path as the convolution of a time-invariant common part and a time-varying variable part. While previous approaches have modeled the common part either using only poles or using only zeros, in this paper we propose to use a common pole-zero model and present an iterative method to compute the common poles and zeros. Using measured acoustic feedback paths from a two-microphone behind-the-ear hearing aid it is shown that the proposed model enables either to increase the modeling accuracy given a fixed number of parameters of the variable part or to reduce the number of parameters of the variable part given a desired accuracy. Henning F. Schepker, Simon Doclo |
ICASSP | 2 |
| 2014 | Single-channel dynamic exemplar-based speech enhancementabstractThis paper proposes an exemplar-based speech enhancement method based on high-resolution STFT magnitude spectrograms, where a selection of the nonnegative training data is used as the dictionary to provide a holistic nonnegative representation of the test data. We discuss how this exemplar-based model ensures that the enhanced speech signal falls on the speech manifold, which improves the quality of the enhanced speech signal. To exploit the temporal continuity, a vector autoregressive model is used to model the activations where the model parameters are learned using a new NMF-based approach. Results from several supervised and semi-supervised speech enhancement experiments indicate that the proposed exemplar-based method outperforms the considered supervised and unsupervised denoising algorithms in terms of both segmental SNR and PESQ at different input SNRs. Index Terms: nonnegative matrix factorization, exemplarbased noise reduction, overcomplete dictionary Nasser Mohammadiha, Simon Doclo |
INTERSPEECH | 2 |
| 2013 | A perceptually constrained channel shortening technique for speech dereverberationabstractThe objective of acoustic multichannel equalization is to design a reshaping filter that reduces reverberation, improves the perceptual speech quality, and is robust to errors in the estimated room impulse responses (RIRs). Although the channel shortening (CS) technique has been shown to be effective in achieving dereverberation, it may fail to preserve the natural shape of an RIR leading to speech quality degradation. Furthermore, CS yields multiple reshaping filters that satisfy its optimization criterion but result in a different perceptual speech quality. In this paper, we propose a robust perceptually constrained channel shortening technique (PeCCS) that resolves the selection ambiguity of CS and leads to joint dereverberation and speech quality preservation. Simulation results for erroneously estimated RIRs show that PeCCS preserves the perceptual speech quality and results in a higher reverberant tail suppression than other state-of-the-art techniques, such as CS and the regularized partial multichannel equalization technique based on the multiple-input/output inverse theorem (P-MINT). Ina Kodrasi, Stefan Goetze, Simon Doclo |
ICASSP | 3 |
| 2013 | Coherence preservation in multi-channel Wiener filtering based noise reduction for binaural hearing aidsabstractBesides noise reduction an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues, i.e. the Interaural Level Difference and the Interaural Time Difference of all sound sources. Recently, extensions of the binaural Multi-channel Wiener filter (MWF) have been presented, which aim to preserve the binaural cues of the residual noise component. However, since these algorithms aim to preserve the Interaural Transfer Function (ITF), they are well-suited only for directional noise sources but not for, e.g. spatially isotropic noise, which can not be fully described by the ITF. In this paper, we present an extension of the binaural MWF, aiming to preserve the Interaural Coherence of the residual noise component. Experimental results using spatially isotropic noise show that the proposed algorithm yields a good preservation of the Interaural Coherence without significantly degrading the output SNR compared to the binaural MWF and the binaural MWF with ITF preservation. Daniel Marquardt, Volker Hohmann, Simon Doclo |
ICASSP | 3 |
| 2013 | Improving speech intelligibility in noise by SII-dependent preprocessing using frequency-dependent amplification and dynamic range compressionabstractIn this contribution, a new preprocessing algorithm to improve speech intelligibility in noise is proposed, which maintains the signal power before and after processing. The proposed Adapt- DRC algorithm consists of two time- And frequency-dependent stages, which are both functions of the estimated SII. The first stage applies a time- And frequency-dependent amplification, while the second stage applies a time- And frequency-dependent dynamic range compression (DRC). Experiments with a competing speaker (CS) and a speech-shaped noise (SSN) show an increase in speech intelligibility for a wide range of SNRs for four different objective measures that are correlated with speech intelligibility. Listening tests conducted within the framework of the Hurricane Challenge with 175 subjects confirm these findings and show improvements of up to 20.5% in intelligibility for SSN and 12.3% for CS. Copyright Henning F. Schepker, Jan Rennies, Simon Doclo |
INTERSPEECH | 3 |
| 2013 | Sound Processing for Better Coding of Monaural and Binaural Cues in Auditory ProsthesesabstractDespite many considerable technical advances in the field of hearing aids and cochlear implants, people using auditory prostheses still have major problems with speech understanding in the presence of interfering sounds and with directional hearing. Both abilities are dependent on sound stream segregation in real-world listening environments. In this paper, two timely and important issues related to sound stream segregation in auditory prostheses are addressed, namely, the coding of monaural and binaural cues. Several state-of-the-art signal processing algorithms used in cochlear implants (CIs) and in hearing aids (HAs) are introduced. A review is given of some recent proposals to improve temporal coding in monaural CIs, and of recent work to improve the transmission of binaural cues in both HAs, CIs, and combined acoustic and electric hearing (bimodal hearing). The ultimate aim is to improve speech and music perception, and, additionally, the preservation of binaural cues to preserve directional hearing. Jan Wouters, Simon Doclo, Raphael Koning, Tom Francart |
Proc. IEEE | 2 |
| 2013 | Regularization for Partial Multichannel Equalization for Speech DereverberationabstractAcoustic multichannel equalization techniques such as the multiple-input/output inverse theorem (MINT), which aim to equalize the room impulse responses (RIRs) between the source and the microphone array, are known to be highly sensitive to RIR estimation errors. To increase robustness, it has been proposed to incorporate regularization in order to decrease the energy of the equalization filters. In addition, more robust partial multichannel equalization techniques such as relaxed multichannel least-squares (RMCLS) and channel shortening (CS) have recently been proposed. In this paper, we propose a partial multichannel equalization technique based on MINT (P-MINT) which aims to shorten the RIR. Furthermore, we investigate the effectiveness of incorporating regularization to further increase the robustness of P-MINT and the aforementioned partial multichannel equalization techniques, i.e., RMCLS and CS. In addition, we introduce an automatic non-intrusive procedure for determining the regularization parameter based on the L-curve. Simulation results using measured RIRs show that incorporating regularization in P-MINT yields a significant performance improvement in the presence of RIR estimation errors, whereas a smaller performance improvement is observed when incorporating regularization in RMCLS and CS. Furthermore, it is shown that the intrusively regularized P-MINT technique outperforms all other investigated intrusively regularized multichannel equalization techniques in terms of perceptual speech quality (PESQ). Finally, it is shown that the automatic non-intrusive regularization parameter in regularized P-MINT leads to a very similar performance as the intrusively determined optimal regularization parameter, making regularized P-MINT a robust, perceptually advantageous, and practically applicable multichannel equalization technique for speech dereverberation. Ina Kodrasi, Stefan Goetze, Simon Doclo |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Robust partial multichannel equalization techniques for speech dereverberationabstractThis paper presents a novel approach for partial multichannel equalization using the multiple-input/output inverse theorem with the first part of one of the estimated channels as the target response (P-MINT). In order to further increase the robustness against channel estimation errors, two extensions are proposed, i.e. the incorporation of a regularization parameter in the inverse filter design and a truncated singular value decomposition approach. Experimental results for speech dereverberation show that the regularized P-MINT method outperforms state-of-the-art techniques such as channel shortening and the relaxed multichannel least-squares method in terms of robustness to channel estimation errors. Ina Kodrasi, Simon Doclo |
ICASSP | 2 |
| 2012 | Binaural cue preservation for hearing aids using multi-channel wiener filter with instantaneous ITF preservationabstractAn important objective of binaural noise reduction algorithms is the preservation of the binaural cues. In this paper an extension of the Multi-channel Wiener filter with binaural cue preservation (MWF-ITF) is presented, where the average noise ITF preservation term is replaced by an instantaneous noise ITF preservation term. This framework in addition allows to impose perfect ITF preservation, leading to a hard-constraint formulation. Experimental results show that the proposed technique yields a better performance in preserving the binaural cues of both the noise component and the speech component compared to the MWF-ITF, without degrading the output SNR. Daniel Marquardt, Volker Hohmann, Simon Doclo |
ICASSP | 3 |
| 2011 | Microphone position optimization for planar superdirective beamformingabstractThe performance of a fixed beamformer highly depends on the position of the microphones in the array. In this paper, different heuristic optimisation approaches for arbitrary planar arrays and an exhaustive search approach for structured array geometries are presented to optimise the microphone positions for a superdirective beamformer, aiming at maximizing the mean directivity index for several steering angles of interest. Through the derivation of an upper bound on the achievable performance, it is shown that the proposed approaches generate configurations with a near-optimal performance. In addition, the theoretical results are validated using real measurements, demonstrating the practical usability of the proposed methods. Ina Kodrasi, Thomas Rohdenburg, Simon Doclo |
ICASSP | 3 |
| 2011 | Analysis of rate constraints for MWF-based noise reduction in acoustic sensor networksabstractIn an acoustic sensor network, consisting of spatially distributed microphone nodes, a significant noise reduction can be achieved using the centralized multi-channel Wiener filter (MWF), requiring all available microphone signals in the entire network. However the limited bandwidth of the communication link typically does not al low to transmit all microphone signals between the different nodes. Recently, a distributed node-specific MWF-based noise reduction scheme has been presented, where each node only transmits a filtered combination of its microphone signals. In this paper, the performance gain of the centralized MWF and the distributed node-specific MWF-based scheme are analyzed as a function of the available band width of the communication link. Toby Christian Lawin-Ore, Simon Doclo |
ICASSP | 2 |
| 2010 | Theoretical Analysis of Binaural Multimicrophone Noise Reduction TechniquesabstractBinaural hearing aids use microphone signals from both left and right hearing aid to generate an output signal for each ear. The microphone signals can be processed by a procedure based on speech distortion weighted multichannel Wiener filtering (SDW-MWF) to achieve significant noise reduction in a speech + noise scenario. In binaural procedures, it is also desirable to preserve binaural cues, in particular the interaural time difference (ITD) and interaural level difference (ILD), which are used to localize sounds. It has been shown in previous work that the binaural SDW-MWF procedure only preserves these binaural cues for the desired speech source, but distorts the noise binaural cues. Two extensions of the binaural SDW-MWF have therefore been proposed to improve the binaural cue preservation, namely the MWF with partial noise estimation (MWF-eta) and MWF with interaural transfer function extension (MWF-ITF). In this paper, the binaural cue preservation of these extensions is analyzed theoretically and tested based on objective performance measures. Both extensions are able to preserve binaural cues for the speech and noise sources, while still achieving significant noise reduction performance. Bram Cornelis, Simon Doclo, Tim Van den Bogaert, Marc Moonen, Jan Wouters |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Reduced-Bandwidth and Distributed MWF-Based Noise Reduction Algorithms for Binaural Hearing AidsabstractIn a binaural hearing aid system, output signals need to be generated for the left and the right ear. Using the binaural multichannel Wiener filter (MWF), which exploits all microphone signals from both hearing aids, a significant reduction of background noise can be achieved. However, due to power and bandwidth limitations of the binaural link, it is typically not possible to transmit all microphone signals between the hearing aids. To limit the amount of transmitted information, this paper presents reduced-bandwidth MWF-based noise reduction algorithms, where a filtered combination of the contralateral microphone signals is transmitted. A first scheme uses a signal-independent beamformer, whereas a second scheme uses the output of a monaural MWF on the contralateral microphone signals and a third scheme involves an iterative distributed MWF (DB-MWF) procedure. It is shown that in the case of a rank-1 speech correlation matrix, corresponding to a single speech source, the DB-MWF procedure converges to the binaural MWF solution. Experimental results compare the noise reduction performance of the reduced-bandwidth algorithms with respect to the benchmark binaural MWF. It is shown that the best performance of the reduced-bandwidth algorithms is obtained by the DB-MWF procedure and that the performance of the DB-MWF procedure approaches quite well the optimal performance of the binaural MWF. Simon Doclo, Marc Moonen, Tim Van den Bogaert, Jan Wouters |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Binaural Cue Preservation for Hearing Aids using an Interaural Transfer Function Multichannel Wiener FilterabstractThis paper describes the binaural cue preservation of a noise reduction algorithm for bilateral hearing aids, namely the multichannel Wiener filter with interaural transfer function extension (MWF-ITF). An extra term is added to the cost function to preserve the binaural cues of both the speech and noise component of a signal at the cost of some noise reduction. This paper combines the theoretical analysis with objective binaural performance measures and a perceptual evaluation. Tim Van den Bogaert, Jan Wouters, Simon Doclo, Marc Moonen |
ICASSP (4) | 3 |
| 2007 | Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction
Simon Doclo, Ann Spriet, Jan Wouters, Marc Moonen |
Speech Commun. | 1 |
| 2007 | Superdirective Beamforming Robust Against Microphone MismatchabstractFixed superdirective beamformers using small-sized microphone arrays are known to be highly sensitive to errors in the assumed microphone array characteristics (gain, phase, position). This paper discusses the design of robust superdirective beamformers by taking into account the statistics of the microphone characteristics. Different design procedures are considered: applying a white noise gain constraint, trading off the mean noise and distortion energy, minimizing the mean deviation from the desired superdirective directivity pattern, and maximizing the mean or the worst case directivity factor. When computational complexity is not an issue, maximizing the mean or the worst case directivity factor is the preferred design procedure. In addition, it is shown how to determine a suitable parameter range for the other design procedures such that both a high directivity and a high level of robustness are obtained Simon Doclo, Marc Moonen |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Superdirective Beamforming Robust Against Microphone MismatchabstractFixed superdirective beamformers using small-size microphone arrays are known to be highly sensitive to errors in the assumed microphone array characteristics. This paper discusses the design of robust superdirective beamformers by taking into account the statistics of the microphone characteristics. Different design procedures are considered: applying a white noise gain constraint, trading off the mean noise and distortion energy, and maximizing the mean or the minimum directivity factor. When computational complexity is not important, maximizing the mean or the minimum directivity factor is the preferred design procedure. In addition, it is shown how to determine a suitable parameter range for the other design procedures Simon Doclo, Marc Moonen |
ICASSP (5) | 1 |
| 2006 | Binaural Multi-Channel Wiener Filtering for Hearing Aids: Preserving Interaural Time and Level DifferencesabstractThis paper presents an extension of the binaural multi-channel Wiener filtering algorithm discussed in T.J. Klasen et al. (2005). The goal of this paper is to preserve both the interaural time difference (ITD) and interaural level difference (ILD) of the speech and noise components. This is done by extending the cost function to incorporate terms for the interaural transfer functions (ITF) of the speech and noise components. Using weights, the emphasis on the preservation of the ITFs can be controlled in addition to the emphasis on noise reduction. Adapting these parameters allows one to preserve the ITFs of the speech and noise component, and therefore ITD and ILD cues, while enhancing the signal-to-noise ratio Thomas J. Klasen, Simon Doclo, Tim Van den Bogaert, Marc Moonen, Jan Wouters |
ICASSP (5) | 2 |
| 2006 | New insights into the noise reduction Wiener filterabstractThe problem of noise reduction has attracted a considerable amount of research attention over the past several decades. Among the numerous techniques that were developed, the optimal Wiener filter can be considered as one of the most fundamental noise reduction approaches, which has been delineated in different forms and adopted in various applications. Although it is not a secret that the Wiener filter may cause some detrimental effects to the speech signal (appreciable or even significant degradation in quality or intelligibility), few efforts have been reported to show the inherent relationship between noise reduction and speech distortion. By defining a speech-distortion index to measure the degree to which the speech signal is deformed and two noise-reduction factors to quantify the amount of noise being attenuated, this paper studies the quantitative performance behavior of the Wiener filter in the context of noise reduction. We show that in the single-channel case the a posteriori signal-to-noise ratio (SNR) (defined after the Wiener filter) is greater than or equal to the a priori SNR (defined before the Wiener filter), indicating that the Wiener filter is always able to achieve noise reduction. However, the amount of noise reduction is in general proportional to the amount of speech degradation. This may seem discouraging as we always expect an algorithm to have maximal noise reduction without much speech distortion. Fortunately, we show that speech distortion can be better managed in three different ways. If we have some a priori knowledge (such as the linear prediction coefficients) of the clean speech signal, this a priori knowledge can be exploited to achieve noise reduction while maintaining a low level of speech distortion. When no a priori knowledge is available, we can still achieve a better control of noise reduction and speech distortion by properly manipulating the Wiener filter, resulting in a suboptimal Wiener filter. In case that we have multiple microphone sensors, the multiple observations of the speech signal can be used to reduce noise with less or even no speech distortion. Jingdong Chen, Jacob Benesty, Yiteng Huang, Simon Doclo |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | On the output SNR of the speech-distortion weighted multichannel Wiener filterabstractIn this letter, we prove that the output signal-to-noise ratio (SNR) after noise reduction with the speech-distortion weighted multichannel Wiener filter is always larger than or equal to the input SNR, for any filter length, for any value of the tradeoff parameter between noise reduction and speech distortion, and for all possible speech and noise correlation matrices. Simon Doclo, Marc Moonen |
IEEE Signal Process. Lett. | 1 |
| 2005 | Multimicrophone noise reduction using recursive GSVD-based optimal filtering with ANC postprocessing stageabstractRecently, a generalized singular value decomposition (GSVD)-based optimal filtering technique has been proposed for enhancing multimicrophone speech signals degraded by additive colored noise. The GSVD-based optimal filtering technique has a better noise reduction performance than standard beamforming techniques provided that the used filter length is large enough. In this paper, it is shown that the same noise reduction performance can be obtained with shorter filter lengths at a lower computational complexity by incorporating the GSVD-based optimal filtering technique in a generalized sidelobe canceller type structure, i.e., by adding an adaptive noise cancellation (ANC) postprocessing stage. Even when using short filter lengths, the total computational complexity is essentially determined by the calculation of the GSVD of a speech and a noise data matrix. It is shown that the complexity can be significantly reduced by using recursive GSVD-updating algorithms and by using subsampling. Simulations have been performed for various acoustic scenarios (different and multiple noise sources and different reverberation conditions), where both the improvement in signal-to-noise ratio and speech distortion have been analyzed. These simulations show that the GSVD-based optimal filtering technique with an ANC postprocessing stage has a better noise reduction performance than standard fixed and adaptive beamforming techniques while introducing an acceptable amount of speech distortion. Simon Doclo, Marc Moonen |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Design of broadband speech beamformers robust against errors in the microphone array characteristicsabstractFixed broadband beamformers for speech applications using small-sized microphone arrays are known to be highly sensitive to errors in the microphone array characteristics. This paper describes two procedures for designing broadband beamformers with an arbitrary spatial directivity pattern, which are robust against gain and phase errors. The first design procedure optimises the mean performance of the broadband beamformer and requires knowledge of the gain and phase probability density functions, whereas the second design procedure optimises the worst-case performance by using a minimax criterion. Simulations with a small-sized microphone array show the performance improvement that can be obtained by using a robust broadband beamformer design procedure. Simon Doclo, Marc Moonen |
ICASSP (5) | 1 |
| 2003 | Design of far-field and near-field broadband beamformers using eigenfilters
Simon Doclo, Marc Moonen |
Signal Process. | 1 |
| 2000 | Combined acoustic echo and noise reduction using GSVD-based optimal filteringabstractThis paper describes two schemes for combining acoustic echo and noise reduction using a GSVD-(generalized singular value decomposition-) based optimal filtering technique. The GSVD-based filtering technique is a signal enhancement technique which has previously been proposed for noise reduction in multi-microphone speech signals. In many speech communication applications however also a far-end echo source is present. Therefore a combined echo and noise reduction scheme is needed. The first scheme combines a standard multi-channel adaptive echo canceller with the GSVD-based noise reduction technique. The second scheme incorporates the far-end echo reference directly into the GSVD-based signal enhancement technique without cancelling the echo in every microphone signal. The two different schemes are compared with regard to performance and computational complexity. Simon Doclo, Marc Moonen, Erik de Clippel |
ICASSP | 1 |
| 1998 | A novel iterative signal enhancement algorithm for noise reduction in speech
Simon Doclo, Ioannis Dologlou, Marc Moonen |
ICSLP | 1 |