Sharon Gannot

dblp:79/5373 · DBLP profile ↗
← Back
125ranked-venue papers
5as first author
21since 2021 · last 2025
0000-0002-2885-170XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 62 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-Microphone Speech Emotion Recognition Using the Hierarchical Token-Semantic Audio Transformer Architecture
abstract
The performance of most emotion recognition systems degrades in real-life situations ("in the wild" scenarios) where the audio is contaminated by reverberation. Our study explores new methods to alleviate the performance degradation of Speech Emotion Recognition (SER) algorithms and develop a more robust system for adverse conditions. We propose processing multi-microphone signals to address these challenges and improve emotion classification accuracy. We adopt a state-of-the-art transformer model, the Hierarchical Token-semantic Audio Transformer (HTS-AT), to handle multi-channel audio inputs. We evaluate two strategies: averaging mel-spectrograms across channels and summing patch-embedded representations. Our multi-microphone model achieves superior performance compared to single-channel baselines when tested on real-world reverberant environments.
Ohad Cohen, Gershon Hazan, Sharon Gannot
ICASSP3
2025 Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
Neta Glazer, David Chernin, Idan Achituve, Sharon Gannot, Ethan Fetaya
INTERSPEECH4
2024 Dataset and Evaluation of Automatic Speech Recognition for Multi-lingual Intent Recognition on Social Robots
abstract
While Automatic Speech Recognition (ASR) systems excel in controlled environments, challenges arise in robot-specific setups due to unique microphone requirements and added noise sources. In this paper, we create a dataset of initiating conversations with brief exchanges in 5 European languages, and we systematically evaluate current state-of-art ASR systems (Vosk, OpenWhisper, Google Speech and NVidia Riva). Besides standard metrics, we also look at two critical downstream tasks for human-robot verbal interaction: intent recognition rate and entity extraction, using the open-source Rasa chatbot. Overall, we found that open-source solutions as Vosk performs competitively with closed-source solutions while running on the edge, on a low compute budget (CPU only).
Antonio Andriella, Raquel Ros, Yoav Ellinson, Sharon Gannot, Séverin Lemaignan
HRI4
2024 Unsupervised Acoustic Scene Mapping Based on Acoustic Features and Dimensionality Reduction
abstract
Classical methods for acoustic scene mapping require the estimation of the time difference of arrival (TDOA) between microphones. Unfortunately, TDOA estimation is very sensitive to reverberation and additive noise. We introduce an unsupervised data-driven approach that exploits the natural structure of the data. Toward this goal, we adapt the recently proposed local conformal autoencoder (LOCA) – an offline deep learning scheme for extracting standardized data coordinates from measurements. Our experimental setup includes a microphone array that measures the transmitted sound source, whose position is unknown, at multiple locations across the acoustic enclosure. We demonstrate that our proposed scheme learns an isometric representation of the microphones’ spatial locations and can perform extrapolation over new and unvisited regions. The performance of our method is evaluated using a series of realistic simulations and compared with a classical approach and other dimensionality-reduction schemes. We further assess reverberation's influence on our framework’s results and show that it demonstrates considerable robustness.
Idan Cohen, Sharon Gannot, Ofir Lindenbaum
ICASSP2
2024 Comparison Of Frequency-Fusion Mechanisms For Binaural Direction-Of-Arrival Estimation For Multiple Speakers
abstract
To estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different frequencies can be combined. According to how the SPS are combined, frequency fusion mechanisms are categorized into narrowband, broadband, or speaker-grouped, where the latter mechanism requires a speaker-wise grouping of frequencies. For a binaural hearing aid setup, in this paper we propose an interaural time difference (ITD)-based speaker-grouped frequency fusion mechanism. By exploiting the DOA dependence of ITDs, frequencies can be grouped according to a common ITD and be used for DOA estimation of the respective speaker. We apply the proposed ITD-based speaker-grouped frequency fusion mechanism for different DOA estimation methods, namely the multiple signal classification, steered response power and a recently published method based on relative transfer function (RTF) vectors. In our experiments, we compare DOA estimation with different fusion mechanisms. For all considered DOA estimation methods, the proposed ITD-based speaker-grouped frequency fusion mechanism results in a higher DOA estimation accuracy compared with the narrowband and broadband fusion mechanisms.
Daniel Fejgin, Elior Hadad, Sharon Gannot, Zbynek Koldovský, Simon Doclo
ICASSP3
2024 LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
abstract
Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for challenging and realistic datasets such as LRS3. In this work, we present LipVoicer, a novel method that generates high-quality speech, even for in-the-wild and rich datasets, by incorporating the text modality. Given a silent video, we first predict the spoken text using a pre-trained lip-reading network. We then condition a diffusion model on the video and use the extracted text through a classifier-guidance mechanism where a pre-trained automatic speech recognition (ASR ) serves as the classifier. LipVoicer outperforms multiple lip-to-speech baselines on LRS2 and LRS3, which are in-the-wild datasets with hundreds of unique speakers in their test set and an unrestricted vocabulary. Moreover, our experiments show that the inclusion of the text modality plays a major role in the intelligibility of the produced speech, readily perceptible while listening, and is empirically reflected in the substantial reduction of the word error rate ( WER ) metric. We demonstrate the effectiveness of LipVoicer through human evaluation, which shows that it produces more natural and synchronized speech signals compared to competing methods. Finally, we created a demo showcasing LipVoicer’s superiority in producing natural, synchronized, and intelligible speech, providing additional evidence of its effectiveness. Project page and code: https://github.com/yochaiye/LipVoicer
Yochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot, Ethan Fetaya
ICLR4
2024 RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
Jacob Bitterman, Daniel Levi, Hilel Hagai Diamandi, Sharon Gannot, Tal Rosenwein
INTERSPEECH4
2024 Efficient Joint Bemforming and Acoustic Echo Cancellation Structure for Conference Call Scenarios
Ofer Schwartz, Sharon Gannot
INTERSPEECH2
2023 Generalized Relative Harmonic Coefficients
abstract
In literature, sound source localization under the far- and near-field scenarios are mostly addressed as independent tasks using different approaches. This causes a tedious task to detect the type of sound-field, whereas in practice there may not be a clear boundary between the far- and near-field soundfield. In contrast, this paper proposes a multi-channel feature denoted generalized relative harmonic coefficients (generalized RHC) in the spherical harmonics domain, which can equally localize both far- and near-field sound source without requiring any adjustments. We derive the analytical expression of this feature and summarize its unique properties, which facilitate two single-source directional-of-arrival estimators: (i) using a full grid search over the directional space; and (ii) a closed-form solution without any grid search. Experimental study in realistic noisy and reverberant environments under both near-field and far-field conditions validates the efficacy of the proposed algorithm.
Yonggang Hu, Sharon Gannot, Thushara D. Abhayapala
ICASSP2
2023 Grad-CAM-Inspired Interpretation of Nearfield Acoustic Holography using Physics-Informed Explainable Neural Network
abstract
The interpretation and explanation of decision-making processes of neural networks are becoming a key factor in the deep learning field. Although several approaches have been presented for classification problems, the application to regression models needs to be further investigated. In this manuscript we propose a Grad-CAM-inspired approach for the visual explanation of neural network architecture for regression problems. We apply this methodology to a recent physics-informed approach for Nearfield Acoustic Holography, called Kirchhoff-Helmholtz-based Convolutional Neural Network (KHCNN) architecture. We focus on the interpretation of KHCNN using vibrating rectangular plates with different boundary conditions and violin top plates with complex shapes. Results highlight the more informative regions of the input that the network exploits to correctly predict the desired output. The devised approach has been validated in terms of NCC and NMSE using the original input and the filtered one coming from the algorithm.
Hagar Kafri, Marco Olivieri, Fabio Antonacci, Mordehay Moradi, Augusto Sarti, Sharon Gannot
ICASSP6
2023 Training-Based Multiple Source Tracking Using Manifold-Learning and Recursive Expectation-Maximization
abstract
In this paper we propose a data-driven approach for multiple speaker tracking in reverberant enclosures. The speakers are uttering, possibly overlapping, speech signals while moving in the environment. The method comprises two stages. The first stage executes a single source localization using semi-supervised learning on multiple manifolds. The second stage, which is unsupervised, uses time-varying maximum likelihood estimation for tracking. The feature vectors, used by both stages, are the relative transfer functions (RTFs), which are known to be related to source positions. The number of sources is assumed to be known while the microphone positions are unknown. In the training stage, a large database of RTFs is given. A small percentage of the data is attributed with exact positions (namely, labelled data) and the rest is assumed to be unlabelled, i.e. the respective position is unknown. Then, a nonlinear, manifold-based, mapping function between the RTFs and the source positions is inferred. Applying this mapping function to all unlabelled RTFs constructs a dense grid of localized sources. In the test phase, this RTF grid serves as the centroids for a Mixture of Gaussians (MoG) model. The MoG parameters are estimated by applying a recursive variant of the expectation-maximization (EM) procedure that relies on the sparsity and intermittency of the speech signals. We present a comprehensive simulation study in various reverberation levels, including static and dynamic scenarios, for both two or three (partially) overlapping speakers. For the dynamic case we provide simulations with several speakers trajectories, including intersecting sources. The proposed scheme outperforms baseline methods that use a simpler propagation model in terms of localization accuracy and tracking capabilities.
Avital Bross, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Closed-Form Single Source Direction-of-Arrival Estimator Using First-Order Relative Harmonic Coefficients
abstract
The relative harmonic coefficients (RHC), recently introduced as a multi-microphone spatial feature, demonstrates promising performance when applied to direction-of-arrival (DOA) estimation. All existing RHC-based DOA estimators suffer from a resolution limitation due to the inherent grid-based search. In contrast, this paper utilizes the first-order RHC to propose a closed-form DOA estimator by deriving a direction vector, which points towards to the desired source direction. Two objective metrics, namely localization accuracy and algorithm complexity, are adopted for the evaluation and comparison with existing RHC-based and intensity based localization approaches, in both simulated and real-life environments.
Yonggang Hu, Sharon Gannot
ICASSP2
2022 Low Resources Online Single-Microphone Speech Enhancement with Harmonic Emphasis
abstract
In this paper, we propose a deep neural network (DNN)-based single-microphone speech enhancement algorithm characterized by a short latency and low computational resources. Many speech enhancement algorithms suffer from low noise reduction capabilities between pitch harmonics, and in severe cases, the harmonic structure may even be lost. Recognizing this drawback, we propose a new weighted loss that emphasizes pitch-dominated frequency bands. For that, we propose a method, applied only at the training stage, to detect these frequency bands. The proposed method is applied to speech signals contaminated by several noise types, and in particular, typical domestic noise drawn from ESC-50 and DE-MAND databases, demonstrating its applicability to ‘stay-at-home’ scenarios.
Nir Raviv, Ofer Schwartz, Sharon Gannot
ICASSP3
2022 A Class of Pareto Optimal Binaural Beamformers
abstract
The objective of binaural multi-microphone speech enhancement algorithms can be viewed as a multi-criteria design problem as there are several requirements to be met. The objective is not only to extract the target speaker without distortion, but also to suppress interfering sources (e.g., competing speakers) and ambient background noise, while preserving the auditory impression of the complete acoustic scene. Such a multi-objective problem (MOP) can be solved using a Pareto frontier, which provides a useful trade-off between the different criteria. In this paper, we propose a unified Pareto optimization framework, which is achieved by defining a generalized mean squared error (MSE) cost function, derived from a MOP. The solution to the multi-criteria problem is grounded on a solid mathematical foundation. The MSE cost function consists of a weighted sum of speech distortion (SD), partial interference reduction (IR), and partial noise reduction (NR) terms with scaling parameters that control the amount of IR and NR. The filter minimizing this generalized cost function, denoted Pareto optimal binaural multichannel Wiener filter (Pareto-BMWF), constitutes a generalization of various binaural MWF-based and binaural MVDR-based beamformers. This solution is optimal for any set of parameters. The improved speech enhancement capabilities are experimentally demonstrated using real-signal recordings when estimation errors are present and the binaural cue preservation capabilities are analyzed.
Elior Hadad, Simon Doclo, Sven Nordholm, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Decoupled Multiple Speaker Direction-of-Arrival Estimator Under Reverberant Environments
abstract
Direction-of-arrival (DOA) estimation for multiple simultaneous speakers in reverberant environments is still one of the challenging tasks in the audio signal processing field. A recent approach addresses this problem using a spherical harmonics domain feature namedrelative harmonic coefficients(RHC). Based on a bin-wise operation across the STFT (short-time Fourier transform) domain, this method detects the direct-path RHC in the first stage, followed by single source localization in the second stage. However, the method is computationally expensive as each STFT bin requires an exhaustive grid search over the two-dimensional (2-D) directional space. In this paper, we propose a significantly more computationally efficient alternative that decouples the azimuth and elevation 2-D search to two separate one-dimensional (1-D) search. The proposed multi-speaker localization algorithm comprises of two main steps, responsible for: (i) achieving a joint direct-path RHC detection and decoupled DOA estimation using 1-D search; and (ii) counting the number of speakers and estimating their DOAs based on the estimates from direct-path dominated STFT bins. Experiments using both simulated and real-life reverberant recordings confirm the significant computational complexity reduction while achieving competitive localization accuracy, compared to the baseline approaches. Although our proposed method performs in an unsupervised manner, it proves to be applicable even under unfavorable acoustic environments with a high reverberation level (e.g.,$T_{60}=1$second).
Yonggang Hu, Prasanga N. Samarasinghe, Sharon Gannot, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Speech Enhancement with Mixture of Deep Experts with Clean Clustering Pre-Training
abstract
In this study we present a mixture of deep experts (MoDE) neural-network architecture for single microphone speech enhancement. Our architecture comprises a set of deep neural networks (DNNs), each of which is an ‘expert’ in a different speech spectral pattern such as phoneme. A gating DNN is responsible for the latent variables which are the weights assigned to each expert’s output given a speech segment. The experts estimate a mask from the noisy input and the final mask is then obtained as a weighted average of the experts’ estimates, with the weights determined by the gating DNN. A soft spectral attenuation, based on the estimated mask, is then applied to enhance the noisy speech signal. As a byproduct, we gain reduction at the complexity in test time. We show that the experts specialization allows better robustness to unfamiliar noise types.1
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
ICASSP3
2021 Evaluation and Comparison of Three Source Direction-of-Arrival Estimators Using Relative Harmonic Coefficients
abstract
A spherical harmonics domain source feature called relative harmonic coefficients (RHC) has recently been applied to address the source direction-of-arrival (DOA) estimation problem. This paper presents a compact evaluation and comparison between two existing RHC based DOA estimators: (i) a method using a full grid search over the two-dimensional (2-D) directional space, (ii) a decoupled estimator which uses one-dimensional (1-D) search to separately localize the source's elevation and azimuth. We also propose a new estimator using a gradient descent search over the 2-D directional grid space. Extensive experiments in both simulated and real-life environments are conducted to examine and analyze the performance of all the underlying DOA estimators. Two objective metrics, including localization accuracy and algorithm complexity, are adopted for an evaluation and comparison between all estimators.
Yonggang Hu, Prasanga N. Samarasinghe, Sharon Gannot, Thushara D. Abhayapala
ICASSP3
2021 Misalignment Recognition in Acoustic Sensor Networks Using a Semi-Supervised Source Estimation Method and Markov Random Fields
abstract
In this paper, we consider the problem of acoustic source localization by acoustic sensor networks (ASNs) using a promising, learning-based technique that adapts to the acoustic environment. In particular, we look at the scenario when a node in the ASN is displaced from its position during training. As the mismatch between the ASN used for learning the localization model and the one after a node displacement leads to erroneous position estimates, a displacement has to be detected and the displaced nodes need to be identified. We propose a method that considers the disparity in position estimates made by leave-one-node-out (LONO) sub-networks and uses a Markov random field (MRF) framework to infer the probability of each LONO position estimate being aligned, misaligned or unreliable while accounting for the noise inherent to the estimator. This probabilistic approach is advantageous over naïve detection methods, as it outputs a normalized value that encapsulates conditional information provided by each LONO sub-network on whether the reading is in misalignment with the overall network. Experimental results confirm that the performance of the proposed method is consistent in identifying compromised nodes in various acoustic conditions.
Gabriel F. Miller, Andreas Brendel, Walter Kellermann, Sharon Gannot
ICASSP4
2021 Online Blind Audio Source Separation Using Recursive Expectation-Maximization
Aviad Eisenberg, Boaz Schwartz, Sharon Gannot
Interspeech3
2021 Scene-Agnostic Multi-Microphone Speech Dereverberation
abstract
Neural networks (NNs) have been widely applied in speech processing tasks, and, in particular, those employing microphone arrays. Nevertheless, most existing NN architectures can only deal with fixed and position-specific microphone arrays. In this paper, we present an NN architecture that can cope with microphone arrays whose number and positions of the microphones are unknown, and demonstrate its applicability in the speech dereverberation task. To this end, our approach harnesses recent advances in deep learning on set-structured data to design an architecture that enhances the reverberant log-spectrum. We use noisy and noiseless versions of a simulated reverberant dataset to test the proposed architecture. Our experiments on the noisy data show that the proposed scene-agnostic setup outperforms a powerful scene-aware framework, sometimes even with fewer microphones. With the noiseless dataset we show that, in most cases, our method outperforms the position-aware network as well as the state-of-the-art weighted linear prediction error (WPE) algorithm.
Yochai Yemini, Ethan Fetaya, Haggai Maron, Sharon Gannot
Interspeech4
2021 Near-Field Superdirectivity: An Analytical Perspective
abstract
The gain achieved by a superdirective beamformer operating in a diffuse noise-field is significantly higher than the gain attainable with conventional delay-and-sum weights. A classical result states that for a compact linear array consisting of N sensors which receives a plane-wave signal from the end-fire direction, the optimal superdirective gain approaches N2. It has been noted that in the near-field regime higher gains can be attained. The gain can increase, in theory, without bound for increasing wavelength or decreasing source-receiver distance. We aim to address the phenomenon of near-field superdirectivity in a comprehensive manner. We derive the optimal performance for the limiting case of an infinitesimal-aperture array receiving a spherical-wave signal. This is done with the aid of a sequence of linear transformations. The resulting gain expression is a polynomial, which depends on the number of sensors employed, the wavelength, and the source-receiver distance. The resulting gain curves are optimal and outperform weights corresponding to other superdirectivity methods. The practical case of a finite-aperture array is discussed. We present conditions for which the gain of such an array would approach that predicted by the theory of the infinitesimal case. The white noise gain (WNG) metric of robustness is shown to increase in the near-field regime.
Dovid Levin, Shmulik Markovich-Golan, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Maximum Likelihood Multi-Speaker Direction of Arrival Estimation Utilizing a Weighted Histogram
abstract
In this contribution, a novel maximum likelihood (ML) based direction of arrival (DOA) estimator for concurrent speakers in a noisy reverberant environment is presented. The DOA estimation task is formulated in the short-time Fourier transform (STFT) in two stages. In the first stage, a single local DOA per time-frequency (TF) bin is selected, using the W-disjoint orthogonality property of the speech signal in the STFT domain. The local DOA is obtained as the maximum of the narrow-band likelihood localization spectrum at each TF bin. In addition, for each local DOA, a confidence measure is calculated, determining the confidence in the local estimate. In the second stage, the wide-band localization spectrum is calculated using a weighted histogram of the local DOA estimates with the confidence measures as weights. Finally, the wide-band DOA estimation is obtained by selecting the peaks in the wide-band localization spectrum. The results of our experimental study demonstrate the benefit of the proposed algorithm in a reverberant environment as compared with the classical steered response power phase transform (SRP-PHAT) algorithm.
Elior Hadad, Sharon Gannot
ICASSP2
2020 Unsupervised Multiple Source Localization Using Relative Harmonic Coefficients
abstract
This paper presents an unsupervised multi-source localization algorithm using a recently introduced feature called the relative harmonic coefficients. We derive a closed-form expression of the feature and briefly summarize its unique properties. We then exploit this feature to develop a single-source frame/bin detector which simplifies the challenging problem of multiple source localization into a single source localization problem. We show that the underlying method is suitable for localization using overlapped, disjoint as well as simultaneous multi-source recordings. Experimental results in both simulated and real-life reverberant environments confirm improved localization accuracy of the proposed method in comparison with the existing state-of-art approach.
Yonggang Hu, Prasanga N. Samarasinghe, Thushara D. Abhayapala, Sharon Gannot
ICASSP4
2020 K-Autoencoders Deep Clustering
abstract
In this study we propose a deep clustering algorithm that extends the k-means algorithm. Each cluster is represented by an autoencoder instead of a single centroid vector. Each data point is associated with the autoencoder which yields the minimal reconstruction error. The optimal clustering is found by learning a set of autoencoders that minimize the global reconstruction mean-square error loss. The network architecture is a simplified version of a previous method that is based on mixture-of-experts. The proposed method is evaluated on standard image corpora and performs on par with state-of-the-art methods which are based on much more complicated network architectures.
Yaniv Opochinsky, Shlomo E. Chazan, Sharon Gannot, Jacob Goldberger
ICASSP3
2020 Low Complexity NLMS for Multiple Loudspeaker Acoustic ECHO Canceller Using Relative Loudspeaker Transfer Functions
abstract
Speech signals captured by a microphone mounted to a smart soundbar or speaker are inherently contaminated by echos. Modern smart devices are usually characterized by low computational capabilities and low memory resources; in these cases, a low-complexity acoustic echo canceller (AEC) may be preferred even though a tolerable degradation in the cancellation occurs. In principle, devices with multiple loudspeakers need an individual AEC for each loudspeaker because the transfer function (TF) from each loudspeaker to the microphone must be estimated. In this paper, we present an normalized least mean square (NLMS) algorithm for a multi-loudspeaker case using relative loudspeaker transfer functions (RLTFs). In each iteration, the RLTFs between each loudspeaker and the reference loudspeaker are estimated first, and then the primary TF between the reference loudspeaker and the microphone. Assuming loudspeakers that are close to each other, the RLTFs can be estimated using fewer coefficients w.r.t. the primary TF, yielding a reduction of 3:4 in computational complexity and 1:2 in memory usage. The algorithm is evaluated using both simulated and real room impulse responses (RIRs) of two loudspeakers with a reverberation time set to 0.3 s and several distances between the loudspeakers.
Ofer Schwartz, Emanuël A. P. Habets, Sharon Gannot
ICASSP3
2020 A Composite DNN Architecture for Speech Enhancement
abstract
In speech enhancement, the use of supervised algorithms in the form of deep neural networks (DNNs) has become tremendously popular in recent years. The target function of the DNN (and the associated estimators) is often either a masking function applied to the noisy spectrum, or the clean log-spectrum. In this work, we show that both separate cost functions are unsuitable for dealing with narrowband noise, and propose a new composite estimator in the log-spectrum domain. The new technique relies on a single DNN that outputs both a masking function and an estimated log-spectrum. Both outputs are used for the composite enhancement. The proposed estimator demonstrates superior performance for speech utterances contaminated by additive narrowband noise, while maintaining the enhancement quality of the baseline algorithms for wideband noise.
Yochai Yemini, Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
ICASSP4
2020 Successive Relative Transfer Function Identification Using Blind Oblique Projection
abstract
Distortionless speech extraction in a reverberant environment can be achieved by applying a beamforming algorithm, provided that the relative transfer functions (RTFs) of the sources and the covariance matrix of the noise are known. In this paper, the challenge of RTF identification in a multi-speaker scenario is addressed. We propose a successive RTF identification (SRI) technique, based on the sole assumption that sources do not become simultaneously active. That is, we address the challenge of estimating the RTF of a specific speech source while assuming that the RTFs of all other active sources in the environment were previously estimated in an earlier stage. The RTF of interest is identified by applying the blind oblique projection (BOP)-SRI technique. When a new speech source is identified, the BOP algorithm is applied. BOP results in a null steering toward the RTF of interest, by means of applying an oblique projection to the microphone measurements. We prove that by artificially increasing the rank of the range of the projection matrix, the RTF of interest can be identified. An experimental study is carried out to evaluate the performance of the BOP-SRI algorithm in various signal to noise ratio (SNR) and signal to interference ratio (SIR) conditions and to demonstrate its effectiveness in speech extraction tasks.
Dani Cherkassky, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Binaural LCMV Beamforming With Partial Noise Estimation
abstract
Besides reducing undesired sources, i.e., interfering sources and background noise, another important objective of a binaural beamforming algorithm is to preserve the spatial impression of the acoustic scene, which can be achieved by preserving the binaural cues of all sound sources. While the binaural minimum variance distortionless response (BMVDR) beamformer provides a good noise reduction performance and preserves the binaural cues of the desired source, it does not allow to control the reduction of the interfering sources and distorts the binaural cues of the interfering sources and the background noise. Hence, several extensions have been proposed. First, the binaural linearly constrained minimum variance (BLCMV) beamformer uses additional constraints, enabling to control the reduction of the interfering sources while preserving their binaural cues. Second, the BMVDR with partial noise estimation (BMVDR-N) mixes the output signals of the BMVDR with the noisy reference microphone signals, enabling to control the binaural cues of the background noise. Aiming at merging the advantages of both extensions, in this paper we propose the BLCMV with partial noise estimation (BLCMV-N). We show that the output signals of the BLCMV-N can be interpreted as a mixture between the noisy reference microphone signals and the output signals of a BLCMV using an adjusted interference scaling parameter. We provide a theoretical comparison between the BMVDR, the BLCMV, the BMVDR-N and the proposed BLCMV-N in terms of noise and interference reduction performance and binaural cue preservation. Experimental results using recorded signals as well as the results of a perceptual listening test show that the BLCMV-N is able to preserve the binaural cues of an interfering source (like the BLCMV), while enabling to trade off between noise reduction performance and binaural cue preservation of the background noise (like the BMVDR-N).
Nico Gößling, Elior Hadad, Sharon Gannot, Simon Doclo
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Semi-Supervised Multiple Source Localization Using Relative Harmonic Coefficients Under Noisy and Reverberant Environments
abstract
This article develops a semi-supervised algorithm to address the challenging multi-source localization problem in a noisy and reverberant environment, using a spherical harmonics domain source feature of the relative harmonic coefficients. We present a comprehensive research of this source feature, including (i) an illustration confirming its sole dependence on the source position, (ii) a feature estimator in the presence of noise, (iii) a feature selector exploiting its inherent directivity over space. Source features at varied spherical harmonic modes, representing unique characterization of the soundfield, are fused by the Multi-Mode Gaussian Process modeling. Based on the unifying model, we then formulate the mapping function revealing the underlying relationship between the source feature(s) and position(s) using a Bayesian inference approach. Another issue of the overlapped components is addressed by a pre-processing technique performing overlapped frame detection, which in turn reduces this challenging problem to a single source localization. It is highlighted that this data-driven method has a strong potential to be implemented in practice because only a limited number of labeled measurements is required. We evaluate this proposed algorithm using simulated recordings between multiple speakers in diverse environments, and extensive results confirm improved performance in comparison with the state-of-art methods. Additional assessments using real-life recordings further prove the effectiveness of the method, even at unfavorable circumstances with severe source overlapping.
Yonggang Hu, Prasanga N. Samarasinghe, Sharon Gannot, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Global and Local Simplex Representations for Multichannel Source Separation
abstract
The problem of blind audio source separation (BASS) in noisy and reverberant conditions is addressed by a novel approach, termed Global and LOcal Simplex Separation (GLOSS), which integrates full- and narrow-band simplex representations. We show that the eigenvectors of the correlation matrix between time frames in a certain frequency band form a simplex that organizes the frames according to the speaker activities in the corresponding band. We propose to build two simplex representations: one global based on a broad frequency band and one local based on a narrow band. In turn, the two representations are combined to determine the dominant speaker in each time-frequency (TF) bin. Using the identified dominating speakers, a spectral mask is computed and is utilized for extracting each of the speakers using spatial beamforming followed by spectral postfiltering. The performance of the proposed algorithm is demonstrated using real-life recordings in various noisy and reverberant conditions.
Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Scoring-Based ML Estimation and CRBs for Reverberation, Speech, and Noise PSDs in a Spatially Homogeneous Noise Field
abstract
Hands-free speech systems are subject to performance degradation due to reverberation and noise. Common methods for enhancing reverberant and noisy speech require the knowledge of the speech, reverberation and noise power spectral densities (PSDs). Most literature on this topic assumes that the noise power spectral density (PSD) matrix is known. However, in many practical acoustic scenarios, the noise PSD is unknown and should be estimated along with the speech and the reverberation PSDs. In this article, the noise is modeled as a spatially homogeneous sound field, with an unknown time-varying PSD multiplied by a known time-invariant spatial coherence matrix. We derive two maximum likelihood estimators (MLEs) for the various PSDs, including the noise: The first is a non-blocking-based estimator, that jointly estimates the PSDs of the speech, reverberation and noise components. The second MLE is a blocking-based estimator, that blocks the speech signal and estimates the reverberation and noise PSDs. Since a closed-form solution does not exist, both estimators iteratively maximize the likelihood using the Fisher scoring method. In order to compare both methods, the corresponding Cramér-Rao Bounds (CRBs) are derived. For both the reverberation and the noise PSDs, it is shown that the non-blocking-based CRB is lower than the blocking-based CRB. Performance evaluation using both simulated and real reverberant and noisy signals, shows that the proposed estimators outperform competing estimators, and greatly reduce the effect of reverberation and noise.
Yaron Laufer, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 ML Estimation and CRBs for Reverberation, Speech, and Noise PSDs in Rank-Deficient Noise Field
abstract
Speech communication systems are prone to performance degradation in reverberant and noisy acoustic environments. Dereverberation and noise reduction algorithms typically require several model parameters, e.g. the speech, reverberation and noise power spectral densities (PSDs). A commonly used assumption is that the noise PSD matrix is known. However, in practical acoustic scenarios, the noise PSD matrix is unknown and should be estimated along with the speech and reverberation PSDs. In this article, we consider the case of rank-deficient noise PSD matrix, which arises when the noise signal consists of multiple directional noise sources, whose number is less than the number of microphones. We derive two closed-form maximum likelihood estimators (MLEs). The first is a non-blocking-based estimator which jointly estimates the speech, reverberation and noise PSDs, and the second is a blocking-based estimator, which first blocks the speech signal and then jointly estimates the reverberation and noise PSDs. Both estimators are analytically compared and analyzed, and mean square errors (MSEs) expressions are derived. Furthermore, Cramér-Rao Bounds (CRBs) on the estimated PSDs are derived. The proposed estimators are examined using both simulation and real reverberant and noisy signals, demonstrating the advantage of the proposed method compared to competing estimators.
Yaron Laufer, Bracha Laufer-Goldshtein, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Simultaneous Tracking and Separation of Multiple Sources Using Factor Graph Model
abstract
In this article, we present an algorithm for direction of arrival (DOA) tracking and separation of multiple speakers with a microphone array using the factor graph statistical model. In our model, the speakers can be located in one of a predefined set of candidate DOAs, and each time-frequency (TF) bin can be associated with a single speaker. Accordingly, by attributing a statistical model to both the DOAs and the associations, as well as to the microphone array observations given these variables, we show that the conditional probability of these variables given the microphone array observations can be modeled as a factor graph. Using the loopy belief propagation (LBP) algorithm, we derive a novel inference scheme which simultaneously estimates both the DOAs and the associations. These estimates are used in turn for separating the sources, by directing a beamformer towards the estimated DOAs, and then applying a TF masking according to the estimated associations. A comprehensive experimental study demonstrates the benefits of the proposed algorithm in both simulated data and real-life measurements recorded in our laboratory.
Koby Weisberg, Bracha Laufer-Goldshtein, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Localization of an Unknown Number of Speakers in Adverse Acoustic Conditions Using Reliability Information and Diarization
abstract
This paper investigates localization of an arbitrary number of simultaneously active speakers in an acoustic enclosure. We propose an algorithm capable of estimating the number of speakers, using reliability information to obtain robust estimation results in adverse acoustic scenarios and estimating individual probability distributions describing the position of each speaker using convex geometry tools. To this end, we start from an established algorithm for localization of acoustic sources based on the EM algorithm. There, the estimation of the number of sources as well as the handling of reverberation has not been addressed sufficiently. We show improvement in the localization of a higher number of sources and in the robustness in adverse conditions including interference from competing speakers, reverberation and noise.
Andreas Brendel, Bracha Laufer-Goldshtein, Sharon Gannot, Ronen Talmon, Walter Kellermann
ICASSP3
2019 An Online Multiple-speaker DOA Tracking Using the CappÉ-Moulines Recursive Expectation-maximization Algorithm
abstract
In this paper, we present a multiple-speaker direction of arrival (DOA) tracking algorithm with a microphone array that utilizes the recursive EM (REM) algorithm proposed by Cappé and Moulines. In our model, all sources can be located in one of a predefined set of candidate DOAs. Accordingly, the received signals from all microphones are modeled as Mixture of Gaussians (MoG) vectors in which each speaker is associated with a corresponding Gaussian. The localization task is then formulated as a maximum likelihood (ML) problem, where the MoG weights and the power spectral density (PSD) of the speakers are the unknown parameters. The REM algorithm is then utilized to estimate the ML parameters in an online manner, facilitating multiple source tracking. By using Fisher-Neyman factorization, the outputs of the minimum variance distortionless response (MVDR)-beamformer (BF) are shown to be sufficient statistics for estimating the parameters of the problem at hand. With that, the terms for the E-step are significantly simplified to a scalar form. An experimental study demonstrates the benefits of the using proposed algorithm in both a simulated data-set and real recordings from the acoustic source localization and tracking (LOCATA) data-set.
Koby Weisberg, Sharon Gannot, Ofer Schwartz
ICASSP2
2019 A Bayesian Hierarchical Model for Speech Enhancement With Time-Varying Audio Channel
abstract
We present a fully Bayesian hierarchical approach for multichannel speech enhancement with time-varying audio channel. Our probabilistic approach relies on a Gaussian prior for the speech signal and a Gamma hyperprior for the speech precision, combined with a multichannel linear-Gaussian state-space model for the acoustic channel. Furthermore, we assume a Wishart prior for the noise precision matrix. We derive a variational expectation-maximization (VEM) algorithm that uses a variant of a multichannel Wiener filter (MCWF) to infer the sound source and a Kalman smoother to infer the acoustic channel. It is further shown that the VEM speech estimator can be recasted as a multichannel minimum variance distortionless response (MVDR) beamformer followed by a single-channel variational postfilter. The proposed algorithm was evaluated using both simulated and real room environments with several noise types and reverberation levels. Both static and dynamic scenarios are considered. In terms of speech quality, it is shown that a significant improvement is obtained with respect to the noisy signal, and that the proposed method outperforms a baseline algorithm. In terms of channel alignment and tracking ability, a superior channel estimate is demonstrated.
Yaron Laufer, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Multichannel Speech Separation and Enhancement Using the Convolutive Transfer Function
abstract
This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, assuming known mixing filters. We propose to perform speech separation and enhancement in the short-time Fourier transform domain using the convolutive transfer function (CTF) approximation. Compared to time-domain filters, the CTF has much less taps. Consequently, it requires less computational cost and sometimes is more robust against the filter perturbations. We propose three methods: 1) for the multisource case, the multichannel inverse filtering method, i.e., the multiple input/output inverse theorem (MINT), is exploited in the CTF domain; 2) a beamforming-like multichannel inverse filtering method applying the single-source MINT and using power minimization, which is suitable whenever the source CTFs are not all known; and 3) a basis pursuit method, where the sources are recovered by minimizing their ℓ1-norm to impose spectral sparsity, while the ℓ2-norm fitting cost between microphone signals and mixing model is constrained to be lower than a tolerance. The noise can be reduced by setting this tolerance at the noise power level. Experiments under various acoustic conditions are carried out to evaluate and compare the three proposed methods. Comparison with four baseline methods-beamforming-based, two time-domain inverse filters, and time-domain Lasso-shows the applicability of the proposed methods.
Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Multichannel Online Dereverberation Based on Spectral Magnitude Inverse Filtering
abstract
This paper addresses the problem of multichannel online dereverberation. The proposed method is carried out in the short-time Fourier transform (STFT) domain, and for each frequency band independently. In the STFT domain, the time-domain room impulse response is approximately represented by the convolutive transfer function (CTF). The multichannel CTFs are adaptively identified based on the cross-relation method, and using the recursive least square criterion. Instead of the complex-valued CTF convolution model, we use a nonnegative convolution model between the STFT magnitude of the source signal and the CTF magnitude, which is just a coarse approximation of the former model, but is shown to be more robust against the CTF perturbations. Based on this nonnegative model, we propose an online STFT magnitude inverse filtering method. The inverse filters of the CTF magnitude are formulated based on the multiple-input/output inverse theorem, and adaptively estimated based on the gradient descent criterion. Finally, the inverse filtering is applied to the STFT magnitude of the microphone signals, obtaining an estimate of the STFT magnitude of the source signal. Experiments regarding both speech enhancement and automatic speech recognition are conducted, which demonstrate that the proposed method can effectively suppress reverberation, even for the difficult case of a moving speaker.
Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming
abstract
In this paper, we present a new control mechanism for LCMV beamforming. Application of the LCMV beamformer to speaker separation tasks requires accurate estimates of its building blocks, e.g. the noise spatial cross-power spectral density (cPSD) matrix and the relative transfer function (RTF) of all sources of interest. An accurate classification of the input frames to various speaker activity patterns can facilitate such an estimation procedure. We propose a DNN-based concurrent speakers detector (CSD) to classify the noisy frames. The CSD, trained in a supervised manner using a DNN, classifies noisy frames into three classes: 1) all speakers are inactive - used for estimating the noise spatial cPSD matrix; 2) a single speaker is active - used for estimating the RTF of the active speaker; and 3) more than one speaker is active - discarded for estimation purposes. Finally, using the estimated blocks, the LCMV beamformer is constructed and applied for extracting the desired speaker from a noisy mixture of speakers.
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
ICASSP3
2018 Multi-View Source Localization Based on Power Ratios
abstract
Despite attracting significant research efforts, the problem of source localization in noisy and reverberant environments remains challenging. Novel learning-based methods attempt to solve the problem by modelling the acoustic environment from the observed data. Typically, appropriate feature vectors are defined, and then used for constructing a model, which maps the extracted features to the corresponding source positions. In this paper, we focus on localizing a source using a distributed network with several arrays of unidirectional microphones. We introduce new feature vectors, which utilize the special characteristic of unidirectional microphones, receiving different parts of the reverberated speech. The new features are computed locally for each array, using the power-ratios between its measured signals, and are used to construct a local model, representing the unique view point of each array. The models of the different arrays, conveying distinct and complementing structures, are merged by a Multi-View Gaussian Process (MVGP), mapping the new features to their corresponding source positions. Based on this unifying model, a Bayesian estimator is derived, exploiting the relations conveyed by the covariance terms of the MVGP. The resulting localizer is shown to be robust to noise and reverberation, utilizing a computationally efficient feature extraction.
Bracha Laufer-Goldshtein, Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP4
2018 A Bayesian Hierarchical Model for Speech Enhancement
abstract
This paper addresses the problem of blind adaptive beamforming using a hierarchical Bayesian model. Our probabilistic approach relies on a Gaussian prior for the speech signal and a Gamma hyperprior for the speech precision, combined with a multichannel linear-Gaussian state-space model for the possibly time-varying acoustic channel. Furthermore, we assume a Gamma prior for the ambient noise precision. We present a variational Expectation-Maximization (VEM) algorithm that employs a variant of multi-channel Wiener filter (MCWF) to estimate the sound source and a Kalman smoother to estimate the acoustic channel of the room. It is further shown that the VEM speech estimator can be decomposed into two stages: A multichannel minimum variance distortionless response (MVDR) beamformer and a subsequent single-channel variational postfilter. The proposed algorithm is evaluated in terms of speech quality, for a static scenario with recorded room impulse responses (RIRs). It is shown that a significant improvement is obtained with respect to the noisy signal, and that the proposed algorithm outperforms a baseline algorithm. In terms of channel alignment, a superior channel estimate is demonstrated compared to the causal Kalman filter.
Yaron Laufer, Sharon Gannot
ICASSP2
2018 Multisource Mint Using Convolutive Transfer Function
abstract
The multichannel inverse filtering method, i.e. multiple input/output inverse theorem (MINT), is widely used. However, it is usually performed in the time domain, and based on the long room impulse responses, thus it has a high computational complexity and a large number of near-common zeros. In this paper, we propose to perform MINT in the short-time Fourier transform (STFT) domain, in which the time-domain filter is approximated by the convolutive transfer function. The oversampled STFT is used to avoid frequency aliasing, which however leads to a common zero region in the subband frequency response due to the frequency response of the STFT window. A new inverse filtering target function concerning the STFT window is proposed to overcome this problem. In addition, unlike most studies using MINT for single source dereverberation, the multisource MINT is proposed for both source separation and dereverberation.
Xiaofei Li 0001, Sharon Gannot, Laurent Girin, Radu Horaud
ICASSP2
2018 DoA Reliability for Distributed Acoustic Tracking
abstract
Distributed acoustic tracking estimates the trajectories of source positions using an acoustic sensor network. As it is often difficult to estimate the source-sensor range from individual nodes, the source positions have to be inferred from the direction-of-arrival (DoA) estimates. Due to reverberation and noise, the sound field becomes increasingly diffuse with increasing source-sensor distance, leading to a decreased Direction of Arrival (DoA)-estimation accuracy. To distinguish between accurate and uncertain DoA estimates, this letter proposes to incorporate the coherent-to-diffuse ratio as a measure of DoA reliability for single-source tracking. It is shown that the source positions, therefore, can be probabilistically triangulated by exploiting the spatial diversity of all nodes.
Christine Evers, Emanuël A. P. Habets, Sharon Gannot, Patrick A. Naylor
IEEE Signal Process. Lett.3
2018 Evaluation and Comparison of Late Reverberation Power Spectral Density Estimators
abstract
Reduction of late reverberation can be achieved using spatio-spectral filters, such as the multichannel Wiener filter. To compute this filter, an estimate of the late reverberation power spectral density (PSD) is required. In recent years, a multitude of late reverberation PSD estimators have been proposed. In this paper, these estimators are categorized into several classes, their relations and differences are discussed, and a comprehensive experimental comparison is provided. To compare their performance, simulations in controlled as well as practical scenarios are conducted. It is shown that a common weakness of spatial coherence-based estimators is their performance in high direct-to-diffuse ratio conditions. To mitigate this problem, a correction method is proposed and evaluated. It is shown that the proposed correction method can decrease the speech distortion without significantly affecting the reverberation reduction.
Sebastian Braun, Adam Kuklasinski, Ofer Schwartz, Oliver Thiergart, Emanuël A. P. Habets, Sharon Gannot, Simon Doclo, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2018 Distributed Expectation-Maximization Algorithm for Speaker Localization in Reverberant Environments
abstract
Localization of acoustic sources has attracted a considerable amount of research attention in recent years. A major obstacle to achieving high localization accuracy is the presence of reverberation, the influence of which obviously increases with the number of active speakers in the room. Human hearing is capable of localizing acoustic sources even in extreme conditions. In this study, we propose to combine a method based on human hearing mechanisms and a modified incremental distributed expectation-maximization (IDEM) algorithm. Rather than using phase difference measurements that are modeled by a mixture of complex-valued Gaussians, as proposed in the original IDEM framework, we propose to use time difference of arrival measurements in multiple subbands and model them by a mixture of real-valued truncated Gaussians. Moreover, we propose to first filter the measurements in order to reduce the effect of the multipath conditions. The proposed method is evaluated using both simulated data and real-life recordings.
Yuval Dorfan, Axel Plinge, Gershon Hazan, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 A Hybrid Approach for Speaker Tracking Based on TDOA and Data-Driven Models
abstract
The problem of speaker tracking in noisy and reverberant enclosures is addressed in this paper. We present a hybrid algorithm, combining traditional tracking schemes with a new learning-based approach. A state-space representation, consisting of a propagation and observation models, is learned from signals measured by several distributed microphone pairs. The proposed representation is based on two data modalities corresponding to high-dimensional acoustic features representing the full reverberant acoustic channels as well as low-dimensional time difference of arrival (TDOA) estimates. The state-space representation is accompanied by a statistical model based on a Gaussian process used to relate the variations of the acoustic channels to the physical variations of the associated source positions, thereby forming a data-driven propagation model for the source movement. In the observation model, the source positions are nonlinearly mapped to the associated TDOA readings. The obtained propagation and observation models establish the basis for employing an extended Kalman filter. The simulation results demonstrate the robustness of the proposed method in noisy and reverberant conditions.
Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Multichannel Identification and Nonnegative Equalization for Dereverberation and Noise Reduction Based on Convolutive Transfer Function
abstract
This paper addresses the problems of blind multichannel identification and equalization for joint speech dereverberation and noise reduction. The time-domain cross-relation method is hardly applicable for blind room impulse response identification due to the near-common zeros of the long impulse responses. We extend the cross-relation method to the short-time Fourier transform (STFT) domain, in which the time-domain impulse response is approximately represented by the convolutive transfer function (CTF) with much less coefficients. For the oversampled STFT, CTFs suffer from the common zeros caused by the nonflat frequency response of the STFT window. To overcome this, we propose to identify CTFs using the STFT framework with oversampled signals and critically sampled CTFs, which is a good tradeoff between the frequency aliasing of the signals and the common zeros problem of CTFs. The identified complex-valued CTFs are not accurate enough for multichannel equalization due to the frequency aliasing of the CTFs. Hence, we only use the CTF magnitudes, which leads to a nonnegative multichannel equalization method based on a nonnegative convolution model between the STFT magnitude of the source signal and the CTF magnitude. Compared with the complex-valued convolution model, this nonnegative convolution model is shown to be more robust against the CTF perturbations. To recover the STFT magnitude of the source signal and to reduce the additive noise, the l2-norm fitting error between the STFT magnitude of the microphone signals and the nonnegative convolution is constrained to be less than a noise power related tolerance. Meanwhile, the l1-norm of the STFT magnitude of the source signal is minimized to impose the sparsity.
Xiaofei Li 0001, Sharon Gannot, Laurent Girin, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Source tracking using moving microphone arrays for robot audition
abstract
Intuitive spoken dialogues are a prerequisite for human-robot interaction. In many practical situations, robots must be able to identify and focus on sources of interest in the presence of interfering speakers. Techniques such as spatial filtering and blind source separation are therefore often used, but rely on accurate knowledge of the source location. In practice, sound emitted in enclosed environments is subject to reverberation and noise. Hence, sound source localization must be robust to both diffuse noise due to late reverberation, as well as spurious detections due to early reflections. For improved robustness against reverberation, this paper proposes a novel approach for sound source tracking that constructively exploits the spatial diversity of a microphone array installed in a moving robot. In previous work, we developed speaker localization approaches using expectation-maximization (EM) approaches and using Bayesian approaches. In this paper we propose to combine the EM and Bayesian approach in one framework for improved robustness against reverberation and noise.
Christine Evers, Yuval Dorfan, Sharon Gannot, Patrick A. Naylor
ICASSP3
2017 Comparison of two binaural beamforming approaches for hearing aids
abstract
Beamforming algorithms in binaural hearing aids are crucial to improve speech understanding in background noise for hearing impaired persons. In this study, we compare and evaluate the performance of two recently proposed minimum variance (MV) beamforming approaches for binaural hearing aids. The binaural linearly constrained MV (BLCMV) beamformer applies linear constraints to maintain the target source and mitigate the interfering sources, taking into account the reverberant nature of sound propagation. The inequality constrained MV (ICMV) beamformer applies inequality constraints to maintain the target source and mitigate the interfering sources, utilizing estimates of the direction of arrivals (DOAs) of the target and interfering sources. The similarities and differences between these two approaches is discussed and the performance of both algorithms is evaluated using simulated data and using real-world recordings, particularly focusing on the robustness to estimation errors of the relative transfer functions (RTFs) and DOAs. The BLCMV achieves a good performance if the RTFs are accurately estimated while the ICMV shows a good robustness to DOA estimation errors.
Elior Hadad, Daniel Marquardt, Wenqiang Pu, Sharon Gannot, Simon Doclo, Zhi-Quan Luo, Ivo Merks, Tao Zhang 0024
ICASSP4
2017 An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixtures
abstract
We present a probabilistic model for joint source separation and diarisation of multichannel convolutive speech mixtures. We build upon the framework of local Gaussian model (LGM) with non-negative matrix factorization (NMF). The diarisation is introduced as a temporal labeling of each source in the mix as active or inactive at the short-term frame level. We devise an EM algorithm in which the source separation process is aided by the diarisation state, since the latter indicates the sources actually present in the mixture. The diarisation state is tracked with a Hidden Markov Model (HMM) with emission probabilities calculated from the estimated source signals. The proposed EM has separation performance comparable with a state-of-the-art LGM NMF method, while outperforming a state-of-the-art speaker diarisation pipeline.
Dionyssos Kounades-Bastian, Laurent Girin, Xavier Alameda-Pineda, Sharon Gannot, Radu Horaud
ICASSP4
2017 Passive Online Geometry Calibration of Acoustic Sensor Networks
abstract
As we are surrounded by an increased number of mobile devices equipped with wireless links and multiple microphones, e.g., smartphones, tablets, laptops, and hearing aids, using them collaboratively for acoustic processing is a promising platform for emerging applications. These devices make up an acoustic sensor network comprised of nodes, i.e., distributed devices equipped with microphone arrays, communication unit, and processing unit. Algorithms for speaker separation and localization using such a network require a precise knowledge of the nodes' locations and orientations. To acquire this knowledge, a recently introduced approach proposed a combined direction of arrival and time difference of arrival (TDoA) target function for offline calibration with dedicated recordings. This letter proposes an extension of this approach to a novel online method with two new features: First, by employing an evolutionary algorithm on incremental measurements, it is online and fast enough for real-time application. Second, by using the sparse spike representation computed in a cochlear model for TDoA estimation, the amount of information shared between the nodes by transmission is reduced, while the accuracy is increased. The proposed approach is able to calibrate an acoustic senor network online during a meeting in a reverberant conference room.
Axel Plinge, Gernot A. Fink, Sharon Gannot
IEEE Signal Process. Lett.3
2017 Blind Synchronization in Wireless Acoustic Sensor Networks
abstract
The challenge of blindly resynchronizing the data acquisition processes in a wireless acoustic sensor network (WASN) is addressed in this paper. The sampling rate offset (SRO) is precisely modeled as a time scaling. The applicability of a wideband correlation processor for estimating the SRO, even in a reverberant and multiple source environment, is presented. An explicit expression for the ambiguity function, which in our case involves time scaling of the received signals, is derived by applying truncated band-limited interpolation. We then propose the recursive band-limited interpolation (RBI) algorithm for recursive SRO estimation. A complete resynchronization scheme utilizing the RBI algorithm, in parallel with the SRO compensation module, is presented. The resulting resynchronization method operates in the time domain in a sequential manner and is, thus, capable of tracking a potentially time-varying SRO. We compared the performance of the proposed RBI algorithm to other available methods in a simulation study. The importance of resynchronization in a beamforming application is demonstrated by both a simulation study and experiments with a real WASN. Finally, we present an experimental study evaluating the expected SRO level between typical data acquisition devices.
Dani Cherkassky, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 A Consolidated Perspective on Multimicrophone Speech Enhancement and Source Separation
abstract
Speech enhancement and separation are core problems in audio signal processing, with commercial applications in devices as diverse as mobile phones, conference call systems, hands-free systems, or hearing aids. In addition, they are crucial preprocessing steps for noise-robust automatic speech and speaker recognition. Many devices now have two to eight microphones. The enhancement and separation capabilities offered by these multichannel interfaces are usually greater than those of single-channel interfaces. Research in speech enhancement and separation has followed two convergent paths, starting with microphone array processing and blind source separation, respectively. These communities are now strongly interrelated and routinely borrow ideas from each other. Yet, a comprehensive overview of the common foundations and the differences between these approaches is lacking at present. In this paper, we propose to fill this gap by analyzing a large number of established and recent techniques according to four transverse axes: 1) the acoustic impulse response model, 2) the spatial filter design criterion, 3) the parameter estimation algorithm, and 4) optional postfiltering. We conclude this overview paper by providing a list of software and data resources and by discussing perspectives and future trends in the field.
Sharon Gannot, Emmanuel Vincent 0001, Shmulik Markovich-Golan, Alexey Ozerov
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Combined LCMV-TRINICON Beamforming for Separating Multiple Speech Sources in Noisy and Reverberant Environments
abstract
The problem of source separation using an array of microphones in reverberant and noisy conditions is addressed. We consider applying the well-known linearly constrained minimum variance (LCMV) beamformer (BF) for extracting individual speakers. Constraints are defined using relative transfer functions (RTFs) for the sources, which are ratios of acoustic transfer functions (ATFs) between any microphone and a reference microphone. The latter are usually estimated by methods that rely on single-talk time segments where only a single source is active and on reliable knowledge of the source activity. Two novel algorithms for estimation of RTFs using the “Triple N” ICA for convolutive mixtures (TRINICON) framework are proposed, not resorting to the usually unavailable source activity pattern. The first algorithm estimates the RTFs of the sources by applying multiple two-channel geometrically constrained (GC) TRINICON units, where approximate direction of arrival information for the sources is utilized for ensuring convergence to the desired solution. The GC-TRINICON is applied to all microphone pairs using a common reference microphone. In the second algorithm, we propose to estimate RTFs iteratively using GC-TRINICON, where instead of using a fixed reference microphone as before, we suggest to use the output signals of LCMV-BFs from the previous iteration as spatially processed references with improved signal-to-interference-and-noise ratio. For both algorithms, a simple detection of noise-only time segments is required for estimating the covariance matrix of noise and interference. We conduct an experimental study in which the performance of the proposed methods is confirmed and compared to corresponding supervised methods.
Shmulik Markovich-Golan, Sharon Gannot, Walter Kellermann
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Semi-Supervised Source Localization on Multiple Manifolds With Distributed Microphones
abstract
The problem of single-source localization with ad hoc microphone networks in noisy and reverberant enclosures is addressed in this paper. A training set is formed by prerecorded measurements collected in advance and consists of a limited number of labelled measurements, attached with corresponding positions, and a larger number of unlabelled measurements from unknown locations. Further information about the enclosure characteristics or the microphone positions is not required. We propose a Bayesian inference approach for estimating a function that maps measurement-based features to the corresponding positions. The signals measured by the microphones represent different viewpoints, which are combined in a unified statistical framework. For this purpose, the mapping function is modelled by a Gaussian process with a covariance function that encapsulates both the connections between pairs of microphones and the relations among the samples in the training set. The parameters of the process are estimated by optimizing a maximum likelihood criterion. In addition, a recursive adaptation mechanism is derived, where the new streaming measurements are used to update the model. Performance is demonstrated for both simulated data and real-life recordings in a variety of reverberation and noise levels.
Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Multiple-Speaker Localization Based on Direct-Path Features and Likelihood Maximization With Spatial Sparsity Regularization
abstract
This paper addresses the problem of multiple-speaker localization in noisy and reverberant environments, using binaural recordings of an acoustic scene. A complex-valued Gaussian mixture model (CGMM) is adopted, whose components correspond to all the possible candidate source locations defined on a grid. After optimizing the CGMM-based objective function, given an observed set of complex-valued binaural features, both the number of sources and their locations are estimated by selecting the CGMM components with the largest weights. An entropy-based penalty term is added to the likelihood to impose sparsity over the set of CGMM component weights. This favors a small number of detected speakers with respect to the large number of initial candidate source locations. In addition, the direct-path relative transfer function (DP-RTF) is used to build robust binaural features. The DP-RTF, recently proposed for single-source localization, encodes interchannel information corresponding to the direct path of sound propagation and is thus robust to reverberations. In this paper, we extend the DP-RTF estimation to the case of multiple sources. In the short-time Fourier transform domain, a consistency test is proposed to check whether a set of consecutive frames is associated with the same source or not. Reliable DP-RTF features are selected from the frames that pass the consistency test to be used for source localization. Experiments carried out using both simulation data and real data recorded with a robotic head confirm the efficiency of the proposed multisource localization method.
Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Multispeaker LCMV Beamformer and Postfilter for Source Separation and Noise Reduction
abstract
The problem of source separation and noise reduction using multiple microphones is addressed. The minimum mean square error (MMSE) estimator for the multispeaker case is derived and a novel decomposition of this estimator is presented. The MMSE estimator is decomposed into two stages: first, a multispeaker linearly constrained minimum variance (LCMV) beamformer (BF); and second, a subsequent multispeaker Wiener postfilter. The first stage separates and enhances the signals of the individual speakers by utilizing the spatial characteristics of the speakers [as manifested by the respective acoustic transfer functions (ATFs)] and the noise power spectral density (PSD) matrix, while the second stage exploits the speakers' PSD matrix to reduce the residual noise at the output of the first stage. The output vector of the multispeaker LCMV BF is proven to be the sufficient statistic for estimating the marginal speech signals in both the classic sense and the Bayesian sense. The log spectral amplitude estimator for the multispeaker case is also derived given the multispeaker LCMV BF outputs. The performance evaluation was conducted using measured ATFs and directional noise with various signal-to-noise ratio levels. It is empirically verified that the multispeaker postfilters are beneficial in terms of signal-to-interference plus noise ratio improvement when compared with the single-speaker postfilter.
Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Cramér-Rao Bound Analysis of Reverberation Level Estimators for Dereverberation and Noise Reduction
abstract
The reverberation power spectral density (PSD) is often required for dereverberation and noise reduction algorithms. In this work, we compare two maximum likelihood (ML) estimators of the reverberation PSD in a noisy environment. In the first estimator, the direct path is first blocked. Then, the ML criterion for estimating the reverberation PSD is stated according to the probability density function of the blocking matrix (BM) outputs. In the second estimator, the speech component is not blocked. Instead, the ML criterion for estimating the speech and reverberation PSD is stated according to the probability density function of the microphone signals. To compare the expected mean square error (MSE) between the two ML estimators of the reverberation PSD, the Cramér-Rao Bounds (CRBs) for the two ML estimators are derived. We show that the CRB for the joint reverberation and speech PSD estimator is lower than the CRB for estimating the reverberation PSD from the BM outputs. Experimental results show that the MSE of the two estimators indeed obeys the CRB curves. Experimental results of multimicrophone dereverberation and noise reduction algorithm show the benefits of using the ML estimators in comparison with another baseline estimators.
Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Two Model-Based EM Algorithms for Blind Source Separation in Noisy Environments
abstract
The problem of blind separation of speech signals in the presence of noise using multiple microphones is addressed. Blind estimation of the acoustic parameters and the individual source signals is carried out by applying the expectation-maximization (EM) algorithm. Two models for the speech signals are used, namely an unknown deterministic signal model and a complex-Gaussian signal model. For the two alternatives, we define a statistical model and develop EM-based algorithms to jointly estimate the acoustic parameters and the speech signals. The resulting algorithms are then compared from both theoretical and performance perspectives. In both cases, the latent data (differently defined for each alternative) are estimated in the E-step, where in the M-step, the two algorithms estimate the acoustic transfer functions of each source and the noise covariance matrix. The algorithms differ in the way the clean speech signals are used in the EM scheme. When the clean signal is assumed deterministic unknown, only the a posteriori probabilities of the presence of each source are estimated in the E-step, whereas their time-frequency coefficients are the parameters that are estimated in the M-step using the minimum variance distortionless response beamformer. If the clean speech signals are modeled as complex Gaussian signals, their power spectral densities are estimated in the E-step using the multichannel Wiener filter output. The proposed algorithms were tested using reverberant noisy mixtures of two speech sources in different reverberation and noise conditions.
Boaz Schwartz, Sharon Gannot, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Extensions of the binaural MWF with interference reduction preserving the binaural cues of the interfering source
abstract
Recently, an extension of the binaural multichannel Wiener filter (BMWF), referred to as BMWF-IRo, was presented in which an interference rejection constraint was added to the BMWF cost function. Although the BMWF-IRo aims to entirely suppress the interfering source, residual interfering sources (as well as unconstrained noise sources) are undesirably perceived as impinging the array from the desired source direction. In this paper, we propose two extensions of the BMWF-IRo that address this issue by preserving the spatial impression of the interfering source. In the first extension, the binaural cues of the interfering source are preserved, while those of the desired source may be slightly distorted. In the second extension, the binaural cues of both the desired and interfering sources are preserved. Simulation results show that the noise reduction performance of both proposed extensions is comparable to the BMWF-, IRo.
Elior Hadad, Daniel Marquardt, Simon Doclo, Sharon Gannot
ICASSP4
2016 An inverse-gamma source variance prior with factorized parameterization for audio source separation
abstract
In this paper we present a new statistical model for the power spectral density (PSD) of an audio signal and its application to multichannel audio source separation (MASS). The source signal is modeled with the local Gaussian model (LGM) and we propose to model its variance with an inverse-Gamma distribution, whose scale parameter is factorized as a rank-1 model. We discuss the interest of this approach and evaluate it in a MASS task with underdetermined convolutive mixtures. For this aim, we derive a variational EM algorithm for parameter estimation and source inference. The proposed model shows a benefit in source separation performance compared to a state-of-the-art LGM NMF-based technique.
Dionyssos Kounades-Bastian, Laurent Girin, Xavier Alameda-Pineda, Sharon Gannot, Radu Horaud
ICASSP4
2016 Manifold-based Bayesian inference for semi-supervised source localization
abstract
Sound source localization is addressed by a novel Bayesian approach using a data-driven geometric model. The goal is to recover the target function that attaches each acoustic sample, formed by the measured signals, with its corresponding position. The estimation is derived by maximizing the posterior probability of the target function, computed on the basis of acoustic samples from known locations (labelled data) as well as acoustic samples from unknown locations (unlabelled data). To form the posterior probability we use a manifold-based prior, which relies on the geometric structure of the manifold from which the acoustic samples are drawn. The proposed method is shown to be analogous to a recently presented semi-supervised localization approach based on manifold regularization. Simulation results demonstrate the robustness of the method in noisy and reverberant environments.
Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot
ICASSP3
2016 Non-stationary noise power spectral density estimation based on regional statistics
abstract
Estimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past and present periodograms in a short-time period. We show that these features are efficient in characterizing the statistical difference between noise PSD and noisy speech PSD. We therefore propose to use these features for estimating the speech presence probability (SPP). The noise PSD is recursively estimated by averaging past spectral power values with a time-varying smoothing parameter controlled by the SPP. The proposed method exhibits good tracking capability for non-stationary noise, even for abruptly increasing noise level.
Xiaofei Li 0001, Laurent Girin, Sharon Gannot, Radu Horaud
ICASSP3
2016 Incorporating relative transfer function preservation into the binaural multi-channel wiener filter for hearing aids
abstract
Besides noise reduction, an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. For the desired speech source and an interfering source, e.g., competing speaker, this can be achieved by preserving their relative transfer functions (RTFs). It has been shown that the binaural multi-channel Wiener filter (MWF) preserves the RTF of the desired speech source, but typically distorts the RTF of the interfering source. To this end, in this paper we propose an extension of the binaural MWF, i.e. the binaural MWF with RTF preservation (MWF-RTF) aiming to preserve the RTF of the interfering source. Analytical expressions for the performance of the binaural MWF and the MWF-RTF in terms of noise reduction and binaural cue preservation are derived, using which their performance is thoroughly compared. Simulation results using binaural behind-the-ear impulse responses measured in a reverberant environment validate the derived analytical expressions, showing that the MWF-RTF yields a better performance than the binaural MWF in terms of the signal-to-interference ratio and binaural cue preservation of the interfering source, while the overall noise reduction performance is slightly degraded.
Daniel Marquardt, Elior Hadad, Sharon Gannot, Simon Doclo
ICASSP3
2016 Joint maximum likelihood estimation of late reverberant and speech power spectral density in noisy environments
abstract
An estimate of the power spectral density (PSD) of the late reverberation is often required by dereverberation algorithms. In this work, we derive a novel multichannel maximum likelihood (ML) estimator for the PSD of the reverberation that can be applied in noisy environments. Since the anechoic speech PSD is usually unknown in advance, it is estimated as well. As a closed-form solution for the maximum likelihood estimator is unavailable, a Newton method for maximizing the ML criterion is derived. Experimental results show that the proposed estimator provides an accurate estimate of the PSD, and outperforms competing estimators. Moreover, when used in a multi-microphone dereverberation and noise reduction algorithm, the best performance in terms of the log-spectral distance is achieved when employing the proposed PSD estimator.
Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets
ICASSP2
2016 Near-field signal acquisition for smartglasses using two acoustic vector-sensors
Dovid Levin, Emanuël A. P. Habets, Sharon Gannot
Speech Commun.3
2016 New Insights into the Kalman Filter Beamformer: Applications to Speech and Robustness
abstract
Statistically optimal spatial processors (also referred to as data-dependent beamformers) are widely-used spatial focusing techniques for desired source extraction. The Kalman filter-based beamformer (KFB) [1] is a recursive Bayesian method for implementing the beamformer. This letter provides new insights into the KFB. Specifically, we adopt the KFB framework to the task of speech extraction. We formalize the KFB with a set of linear constraints and present its equivalence to the linearly constrained minimum power (LCMP) beamformer. We further show that the optimal output power, required for implementing the KFB, is merely controlling the white noise gain (WNG) of the beamformer. We also show, that in static scenarios, the adaptation rule of the KFB reduces to the simpler affine projection algorithm (APA). The analytically derived results are verified and exemplified by a simulation study.
Dani Cherkassky, Sharon Gannot
IEEE Signal Process. Lett.2
2016 A Hybrid Approach for Speech Enhancement Using MoG Model and Neural Network Phoneme Classifier
abstract
In this paper, we present a single-microphone speech enhancement algorithm. A hybrid approach is proposed merging the generative mixture of Gaussians (MoG) model and the discriminative deep neural network (DNN). The proposed algorithm is executed in two phases, the training phase, which does not recur, and the test phase. First, the noise-free speech log-power spectral density is modeled as an MoG, representing the phoneme-based diversity in the speech signal. A DNN is then trained with phoneme labeled database of clean speech signals for phoneme classification with mel-frequency cepstral coefficients as the input features. In the test phase, a noisy utterance of an untrained speech is processed. Given the phoneme classification results of the noisy speech utterance, a speech presence probability (SPP) is obtained using both the generative and discriminative models. SPP-controlled attenuation is then applied to the noisy speech while simultaneously, the noise estimate is updated. The discriminative DNN maintains the continuity of the speech and the generative phoneme-based MoG preserves the speech spectral structure. Extensive experimental study using real speech and noise signals is provided. We also compare the proposed algorithm with alternative speech enhancement algorithms. We show that we obtain a significant improvement over previous methods in terms of speech quality measures. Finally, we analyze the contribution of all components of the proposed algorithm indicating their combined importance.
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 The Binaural LCMV Beamformer and its Performance Analysis
abstract
The recently proposed binaural linearly constrained minimum variance (BLCMV) beamformer is an extension of the well-known binaural minimum variance distortionless response (MVDR) beamformer, imposing constraints for both the desired and the interfering sources. Besides its capabilities to reduce interference and noise, it also enables to preserve the binaural cues of both the desired and interfering sources, hence making it particularly suitable for binaural hearing aid applications. In this paper, a theoretical analysis of the BLCMV beamformer is presented. In order to gain insights into the performance of the BLCMV beamformer, several decompositions are introduced that reveal its capabilities in terms of interference and noise reduction, while controlling the binaural cues of the desired and the interfering sources. When setting the parameters of the BLCMV beamformer, various considerations need to be taken into account, e.g. based on the amount of interference and noise reduction and the presence of estimation errors of the required relative transfer functions (RTFs). Analytical expressions for the performance of the BLCMV beamformer in terms of noise reduction, interference reduction, and cue preservation are derived. Comprehensive simulation experiments, using measured acoustic transfer functions as well as real recordings on binaural hearing aids, demonstrate the capabilities of the BLCMV beamformer in various noise environments.
Elior Hadad, Simon Doclo, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 A Variational EM Algorithm for the Separation of Time-Varying Convolutive Audio Mixtures
abstract
This paper addresses the problem of separating audio sources from time-varying convolutive mixtures. We propose a probabilistic framework based on the local complex-Gaussian model combined with non-negative matrix factorization. The time-varying mixing filters are modeled by a continuous temporal stochastic process. We present a variational expectation-maximization (VEM) algorithm that employs a Kalman smoother to estimate the time-varying mixing matrix, and that jointly estimate the source parameters. The sound sources are then separated by Wiener filters constructed with the estimators provided by the VEM algorithm. Extensive experiments on simulated data show that the proposed method outperforms a blockwise version of a state-of-the-art baseline method.
Dionyssos Kounades-Bastian, Laurent Girin, Xavier Alameda-Pineda, Sharon Gannot, Radu Horaud
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Semi-Supervised Sound Source Localization Based on Manifold Regularization
abstract
Conventional speaker localization algorithms, based merely on the received microphone signals, are often sensitive to adverse conditions, such as: high reverberation or low signal-to-noise ratio (SNR). In some scenarios, e.g., in meeting rooms or cars, it can be assumed that the source position is confined to a predefined area, and the acoustic parameters of the environment are approximately fixed. Such scenarios give rise to the assumption that the acoustic samples from the region of interest have a distinct geometrical structure. In this paper, we show that the high-dimensional acoustic samples indeed lie on a low-dimensional manifold and can be embedded into a low-dimensional space. Motivated by this result, we propose a semi-supervised source localization algorithm based on two-microphone measurements, which recovers the inverse mapping between the acoustic samples and their corresponding locations. The idea is to use an optimization framework based on manifold regularization, that involves smoothness constraints of possible solutions with respect to the manifold. The proposed algorithm, termed manifold regularization for localization, is adapted while new unlabelled measurements (from unknown source locations) are accumulated during runtime. Experimental results show superior localization performance when compared with a recently presented algorithm based on a manifold learning approach and with the generalized cross-correlation algorithm as a baseline. The algorithm achieves 2° accuracy in typical noisy and reverberant environments (reverberation time between 200 and 800 ms and SNR between 5 and 20 dB).
Bracha Laufer-Goldshtein, Ronen Talmon, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Estimation of the Direct-Path Relative Transfer Function for Supervised Sound-Source Localization
abstract
This paper addresses the problem of sound-source localization of a single speech source in noisy and reverberant environments. For a given binaural microphone setup, the binaural response corresponding to the direct-path propagation of a single source is a function of the source direction. In practice, this response is contaminated by noise and reverberations. The direct-path relative transfer function (DP-RTF) is defined as the ratio between the direct-path acoustic transfer function of the two channels. We propose a method to estimate the DP-RTF from the noisy and reverberant microphone signals in the short-time Fourier transform (STFT) domain. First, the convolutive transfer function approximation is adopted to accurately represent the impulse response of the sensors in the STFT domain. Second, the DP-RTF is estimated by using the auto- and cross-power spectral densities at each frequency and over multiple frames. In the presence of stationary noise, an interframe spectral subtraction algorithm is proposed, which enables to achieve the estimation of noise-free auto- and cross-power spectral densities. Finally, the estimated DP-RTFs are concatenated across frequencies and used as a feature vector for the localization of speech source. Experiments with both simulated and real data show that the proposed localization method performs well, even under severe adverse acoustic conditions, and outperforms state-of-the-art localization methods under most of the acoustic conditions.
Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 An Expectation-Maximization Algorithm for Multimicrophone Speech Dereverberation and Noise Reduction With Coherence Matrix Estimation
abstract
In speech communication systems, the microphone signals are degraded by reverberation and ambient noise. The reverberant speech can be separated into two components, namely, an early speech component that consists of the direct path and some early reflections and a late reverberant component that consists of all late reflections. In this paper, a novel algorithm to simultaneously suppress early reflections, late reverberation, and ambient noise is presented. The expectation-maximization (EM) algorithm is used to estimate the signals and spatial parameters of the early speech component and the late reverberation components. As a result, a spatially filtered version of the early speech component is estimated in the E-step. The power spectral density (PSD) of the anechoic speech, the relative early transfer functions, and the PSD matrix of the late reverberation are estimated in the M-step of the EM algorithm. The algorithm is evaluated using real room impulse response recorded in our acoustic lab with a reverberation time set to 0.36 s and 0.61 s and several signal-to-noise ratio levels. It is shown that significant improvement is obtained and that the proposed algorithm outperforms baseline single-channel and multichannel dereverberation algorithms, as well as a state-of-the-art multichannel dereverberation algorithm.
Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whitening method
abstract
Microphone array processing utilize spatial separation between the desired speaker and interference signal for speech enhancement. The transfer functions (TFs) relating the speaker component at a reference microphone with all other microphones, denoted as the relative TFs (RTFs), play an important role in beamforming design criteria such as minimum variance distortionless response (MVDR) and speech distortion weighted multichannel Wiener filter (SDW-MWF). Two common methods for estimating the RTF are surveyed here, namely, the covariance subtraction (CS) and the covariance whitening (CW) methods. We analyze the performance of the CS method theoretically and empirically validate the results of the analysis through extensive simulations. Furthermore, empirically comparing the methods performances in various scenarios evidently shows thats the CW method outperforms the CS method.
Shmulik Markovich-Golan, Sharon Gannot
ICASSP2
2015 Binaural multichannel Wiener filter with directional interference rejection
abstract
In this paper we consider an acoustic scenario with a desired source and a directional interference picked up by hearing devices in a noisy and reverberant environment. We present an extension of the binaural multichannel Wiener filter (BMWF), by adding an interference rejection constraint to its cost function, in order to combine the advantages of spatial and spectral filtering while mitigating directional interferences. We prove that this algorithm can be decomposed into the binaural linearly constrained minimum variance (BLCMV) algorithm followed by a single channel Wiener post-filter. The proposed algorithm yields improved interference rejection capabilities, as compared with the BMWF. Moreover, by utilizing the spectral information on the sources, it is demonstrating better SNR measures, as compared with the BLCMV.
Elior Hadad, Daniel Marquardt, Simon Doclo, Sharon Gannot
ICASSP4
2015 Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtraction
abstract
This paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresponding to speech-plus-noise activity and noise-only. Then, the subtraction of two segmental PSD matrices leads to an almost noise-free PSD matrix by reducing the stationary noise component and preserving non-stationary speech component. This noise-free PSD matrix is used for single speaker RTF identification by eigenvalue decomposition. Experiments are performed in the context of sound source localization to evaluate the efficiency of the proposed method.
Xiaofei Li 0001, Laurent Girin, Radu Horaud, Sharon Gannot
ICASSP4
2015 Nested generalized sidelobe canceller for joint dereverberation and noise reduction
abstract
Speech signal is often contaminated by both room reverberation and ambient noise. In this contribution, we propose a nested generalized sidelobe canceller (GSC) beamforming structure, comprising an inner and an outer GSC beamformers (BFs), that decouple the speech dereverberation and the noise reduction operations. The BFs are implemented in the short-time Fourier transform (STFT) domain. Two alternative reverberation models are adopted. In the first, used in the inner GSC, reverberation is assumed to comprise a coherent early component and a late reverberant component. In the second, used in the outer GSC, the influence of the entire acoustic transfer function (ATF) is modeled as a convolution along the frame index in each frequency. Unlike other BF designs for this problem that must be updated in each time-frame, the proposed BF is time-invariant in static scenarios. Experiments with both simulated and recorded environments verify the effectiveness of the proposed structure.
Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets
ICASSP2
2015 Special issue on wireless acoustic sensor networks and ad hoc microphone arrays
Alexander Bertrand, Simon Doclo, Sharon Gannot, Nobutaka Ono, Toon van Waterschoot
Signal Process.3
2015 Optimal distributed minimum-variance beamforming approaches for speech enhancement in wireless acoustic sensor networks
Shmulik Markovich-Golan, Alexander Bertrand, Marc Moonen, Sharon Gannot
Signal Process.4
2015 On the Average Directivity Factor Attainable With a Beamformer Incorporating Null Constraints
abstract
The directivity factor (DF) of a beamformer describes its spatial selectivity and ability to suppress diffuse noise which arrives from all directions. For a given array constellation, it is possible to select beamforming weights which maximize the DF for a particular look-direction, while enforcing nulls for a set of undesired directions. In general, the resulting DF is dependent upon the specific look- and null directions. Using the same array, one may apply a different set of weights designed for any other feasible set of look- and null directions. In this contribution, we show that when the optimal DF is averaged over all look directions, the result equals the number of sensors minus the number of null constraints. This result holds regardless of the positions and spatial responses of the individual sensors and regardless of the null directions. The result generalizes to more complex wave-propagation domains (e.g., reverberation).
Dovid Levin, Emanuël A. P. Habets, Sharon Gannot
IEEE Signal Process. Lett.3
2015 Tree-Based Recursive Expectation-Maximization Algorithm for Localization of Acoustic Sources
abstract
The problem of distributed localization for ad hoc wireless acoustic sensor networks (WASNs) is addressed in this paper. WASNs are characterized by low computational resources in each node and by limited connectivity between the nodes. Novel bi-directional tree-based distributed estimation–maximization (DEM) algorithms are proposed to circumvent these inherent limitations. We show that the proposed algorithms are capable of localizing static acoustic sources in reverberant enclosures without a priori information on the number of sources. Unlike serial estimation procedures (like ring-based algorithms), the new algorithms enable simultaneous computations in the nodes and exhibit greater robustness to communication failures. Specifically, the recursive distributed EM (RDEM) variant is better suited to online applications due to its recursive nature. Furthermore, the RDEM outperforms the other proposed variants in terms of convergence speed and simplicity. Performance is demonstrated by an extensive experimental study consisting of both simulated and actual environments.
Yuval Dorfan, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Theoretical Analysis of Binaural Transfer Function MVDR Beamformers with Interference Cue Preservation Constraints
abstract
The objective of binaural noise reduction algorithms is not only to selectively extract the desired speaker and to suppress interfering sources (e.g., competing speakers) and ambient background noise, but also to preserve the auditory impression of the complete acoustic scene. For directional sources this can be achieved by preserving the relative transfer function (RTF) which is defined as the ratio of the acoustical transfer functions relating the source and the two ears and corresponds to the binaural cues. In this paper, we theoretically analyze the performance of three algorithms that are based on the binaural minimum variance distortionless response (BMVDR) beamformer, and hence, process the desired source without distortion. The BMVDR beamformer preserves the binaural cues of the desired source but distorts the binaural cues of the interfering source. By adding an interference reduction (IR) constraint, the recently proposed BMVDR-IR beamformer is able to preserve the binaural cues of both the desired source and the interfering source. We further propose a novel algorithm for preserving the binaural cues of both the desired source and the interfering source by adding a constraint preserving the RTF of the interfering source, which will be referred to as the BMVDR-RTF beamformer. We analytically evaluate the performance in terms of binaural signal-to-interference-and-noise ratio (SINR), signal-to-interference ratio (SIR), and signal-to-noise ratio (SNR) of the three considered beamformers. It can be shown that the BMVDR-RTF beamformer outperforms the BMVDR-IR beamformer in terms of SINR and outperforms the BMVDR beamformer in terms of SIR. Among all beamformers which are distortionless with respect to the desired source and preserve the binaural cues of the interfering source, the newly proposed BMVDR-RTF beamformer is optimal in terms of SINR. Simulations using acoustic transfer functions measured on a binaural hearing aid validate our theoretical results.
Elior Hadad, Daniel Marquardt, Simon Doclo, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Spatial Source Subtraction Based on Incomplete Measurements of Relative Transfer Function
abstract
Relative impulse responses between microphones are usually long and dense due to the reverberant acoustic environment. Estimating them from short and noisy recordings poses a long-standing challenge of audio signal processing. In this paper, we apply a novel strategy based on ideas of compressed sensing. Relative transfer function (RTF) corresponding to the relative impulse response can often be estimated accurately from noisy data but only for certain frequencies. This means that often only an incomplete measurement of the RTF is available. A complete RTF estimate can be obtained through finding its sparsest representation in the time-domain: that is, through computing the sparsest among the corresponding relative impulse responses. Based on this approach, we propose to estimate the RTF from noisy data in three steps. First, the RTF is estimated using any conventional method such as the nonstationarity-based estimator by Gannotor through blind source separation. Second, frequencies are determined for which the RTF estimate appears to be accurate. Third, the RTF is reconstructed through solving a weighted${\ell _1}$convex program, which we propose to solve via a computationally efficient variant of the SpaRSA (Sparse Reconstruction by Separable Approximation) algorithm. An extensive experimental study with real-world recordings has been conducted. It has been shown that the proposed method is capable of improving many conventional estimators used as the first step in most situations.
Zbynek Koldovský, Jirí Málek, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Theoretical Analysis of Linearly Constrained Multi-Channel Wiener Filtering Algorithms for Combined Noise Reduction and Binaural Cue Preservation in Binaural Hearing Aids
abstract
Besides noise reduction, an important objective of binaural speech enhancement algorithms is the preservation of the binaural cues of all sound sources. For the desired speech source and the interfering sources, e.g., competing speakers, this can be achieved by preserving their relative transfer functions (RTFs). It has been shown that the binaural multi-channel Wiener filter (MWF) preserves the RTF of the desired speech source, but typically distorts the RTF of the interfering sources. To this end, in this paper we propose two extensions of the binaural MWF, i.e., the binaural MWF with RTF preservation (MWF-RTF) aiming to preserve the RTF of the interfering source and the binaural MWF with interference rejection (MWF-IR) aiming to completely suppress the interfering source. Analytical expressions for the performance of the binaural MWF, MWF-RTF and MWF-IR in terms of noise reduction, speech distortion and binaural cue preservation are derived, showing that the proposed extensions yield a better performance in terms of the signal-to-interference ratio and preservation of the binaural cues of the directional interference, while the overall noise reduction performance is degraded compared to the binaural MWF. Simulation results using binaural behind-the-ear impulse responses measured in a reverberant environment validate the derived analytical expressions for the theoretically achievable performance of the binaural MWF, MWF-RTF, and MWF-IR, showing that the performance highly depends on the position of the interfering source and the number of microphones. Furthermore, the simulation results show that the MWF-RTF yields a very similar overall noise reduction performance as the binaural MWF, while preserving the binaural cues of both the speech and the interfering source.
Daniel Marquardt, Elior Hadad, Sharon Gannot, Simon Doclo
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Multi-Microphone Speech Dereverberation and Noise Reduction Using Relative Early Transfer Functions
abstract
In speech communication systems, the microphone signals are degraded by reverberation and ambient noise. The reverberant speech can be separated into two components, namely, an early speech component that includes the direct path and some early reflections, and a late reverberant component that includes all the late reflections. In this paper, a novel algorithm to simultaneously suppress early reflections, late reverberation and ambient noise is presented. A multi-microphone minimum mean square error estimator is used to obtain a spatially filtered version of the early speech component. The estimator constructed as a minimum variance distortionless response (MVDR) beamformer (BF) followed by a postfilter (PF). Three unique design features characterize the proposed method. First, the MVDR BF is implemented in a special structure, named the nonorthogonal generalized sidelobe canceller (NO-GSC). Compared with the more conventional orthogonal GSC structure, the new structure allows for a simpler implementation of the GSC blocks for various MVDR constraints. Second, In contrast to earlier works, RETFs are used in the MVDR criterion rather than either the entire RTFs or only the direct-path of the desired speech signal. An estimator of the RETFs is proposed as well. Third, the late reverberation and noise are processed by both the beamforming stage and the PF stage. Since the relative power of the noise and the late reverberation varies with the frame index, a computationally efficient method for the required matrix inversion is proposed to circumvent the cumbersome mathematical operation. The algorithm was evaluated and compared with two alternative multichannel algorithms and one single-channel algorithm using simulated data and data recorded in a room with a reverberation time of 0.5 s for various source-microphone array distances (1-4 m) and several signal-to-noise levels. The processed signals were tested using two commonly used objective measures, namely perceptual evaluation of speech quality and log-spectral distance. As an additional objective measure, the improvement in word accuracy percentage of an acoustic speech recognition system is also demonstrated.
Ofer Schwartz, Sharon Gannot, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Online Speech Dereverberation Using Kalman Filter and EM Algorithm
abstract
Speech signals recorded in a room are commonly degraded by reverberation. In most cases, both the speech signal and the acoustic system of the room are unknown and time-varying. In this paper, a scenario with a single desired sound source and slowly time-varying and spatially-white noise is considered, and a multi-microphone algorithm that simultaneously estimates the clean speech signal and the time-varying acoustic system is proposed. The recursive expectation-maximization scheme is employed to obtain both the clean speech signal and the acoustic system in an online manner. In the expectation step, the Kalman filter is applied to extract a new sample of the clean signal, and in the maximization step, the system estimate is updated according to the output of the Kalman filter. Experimental results show that the proposed method is able to significantly reduce reverberation and increase the speech quality. Moreover, the tracking ability of the algorithm was validated in practical scenarios using human speakers moving in a natural manner.
Boaz Schwartz, Sharon Gannot, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Multichannel Wiener filter performance analysis in presence of mismodeling
abstract
A randomly positioned microphone array is considered in this work. In many applications, the locations of the array elements are known up to a certain degree of random mismatch. We derive a novel statistical model for performance analysis of the multi-channel Wiener filter (MWF) beamformer under random mismatch in sensors location. We consider the scenario of one desired source and one interfering source arriving from the far-field and impinging on a linear array. A theoretical model for predicting the MWF mean squared error (MSE) for a given variation in sensors location is developed and verified by simulations. It is postulated that the probability density function (p.d.f) of the MSE of the MWF obeys Γ distribution. This claim is verified empirically by simulations.
Dani Cherkassky, Sharon Gannot
ICASSP2
2014 Speaker Tracking Using Recursive EM Algorithms
abstract
The problem of localizing and tracking a known number of concurrent speakers in noisy and reverberant enclosures is addressed in this paper. We formulate the localization task as a maximum likelihood (ML) parameter estimation problem, and solve it by utilizing the expectation-maximization (EM) procedure. For the tracking scenario, we propose to adapt two recursive EM (REM) variants. The first, based on Titterington's scheme, is a Newton-based recursion. In this work we also extend Titterington's method to deal with constrained maximization, encountered in the problem at hand. The second is based on Cappé and Moulines' scheme. We discuss the similarities and dissimilarities of these two variants and show their applicability to the tracking problem by a simulated experimental study.
Ofer Schwartz, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Robust beamforming using sensors with nonidentical directivity patterns
abstract
The optimal weights for a beamformer that provide maximum directivity, are often found to be severely lacking in terms of robustness. Although an ideal implementation of the beamformer with these weights provides high directivity, minor perturbations of the weights or of sensor placement cause severe degradation. Therefore, a robustness constraint is often imposed during the beamformer's design stage. The classical method of diagonal loading is commonly used for this purpose. There are known results in this field which pertain to an array consisting of sensors with identical directivity-patterns and orientations. We extend these results to account for sensors with nonidentical directivity patterns, and sensors which share placement errors. We show that in such cases, modification of the classical loading scheme to incorporate nonidentical diagonal elements and off-diagonal elements is beneficial.
Dovid Levin, Emanuël A. P. Habets, Sharon Gannot
ICASSP3
2013 A Generalized Theorem on the Average Array Directivity Factor
abstract
The beampattern of an array consisting of$N$elements is determined by the beampatterns of the individual elements, their placement, and the weights assigned to them. For each look direction, it is possible to design weights that maximize the array directivity factor (DF). For the case of an array of omnidirectional elements using optimal weights, it has been shown that the average DF over all look directions equals the number of elements. The validity of this theorem is not dependent on array geometry. We generalize this theorem by means of an alternative proof. The chief contributions of this letter are a) a compact and direct proof, b) generalization to arrays containing directional elements (such as cardioids and dipoles), and c) generalization to arbitrary wave propagation models. A discussion of the theorem's ramifications on array processing is provided.
Dovid Levin, Emanuël A. P. Habets, Sharon Gannot
IEEE Signal Process. Lett.3
2013 Distributed Multiple Constraints Generalized Sidelobe Canceler for Fully Connected Wireless Acoustic Sensor Networks
abstract
This paper proposes a distributed multiple constraints generalized sidelobe canceler (GSC) for speech enhancement in anN-node fully connected wireless acoustic sensor network (WASN) comprisingMmicrophones. Our algorithm is designed to operate in reverberant environments with constrained speakers (including both desired and competing speakers). Rather than broadcastingMmicrophone signals, a significant communication bandwidth reduction is obtained by performing local beamforming at the nodes, and utilizing only transmission channels. Each node processes its own microphone signals together with the N + P transmitted signals. The GSC-form implementation, by separating the constraints and the minimization, enables the adaptation of the BF during speech-absent time segments, and relaxes the requirement of other distributed LCMV based algorithms to re-estimate the sources RTFs after each iteration. We provide a full convergence proof of the proposed structure to the centralized GSC-beamformer (BF). An extensive experimental study of both narrowband and (wideband) speech signals verifíes the theoretical analysis.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.2
2013 Performance of the SDW-MWF With Randomly Located Microphones in a Reverberant Enclosure
abstract
Beamforming with wireless acoustic sensor networks (WASNs) has recently drawn the attention of the research community. As the number of microphones grows it is difficult, and in some applications impossible, to determine their layout beforehand. A common practice in analyzing the expected performance is to utilize statistical considerations. In the current contribution, we consider applying the speech distortion weighted multi-channel Wiener filter (SDW-MWF) to enhance a desired source propagating in a reverberant enclosure where the microphones are randomly located with a uniform distribution. Two noise fields are considered, namely, multiple coherent interference signals and a diffuse sound field. Utilizing the statistics of the acoustic transfer function (ATF), we derive a statistical model for two important criteria of the beamformer (BF): the signal to interference ratio (SIR), and the white noise gain. Moreover, we propose reliability functions, which determine the probability of the SIR and white noise gain to exceed a predefined level. We verify the proposed model with an extensive simulative study.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.2
2013 Single-Channel Transient Interference Suppression With Diffusion Maps
abstract
A transient is an abrupt or impulsive sound followed by decaying oscillations, e.g., keyboard typing and door knocking. Such sounds often arise as interference in everyday applications, e.g., hearing aids, hands-free accessories, mobile phones, and conference-room devices. In this paper, we present an algorithm for single-channel transient interference suppression. The main component of the proposed algorithm is the estimation of the spectral variance of the interference. We propose a statistical model of the transient interference and combine it with non-local filtering. We exploit the unique spectral structure of the transients along with their impulsive temporal nature to distinct them from speech. A particular attention is given to handling both short- and long-duration transients. Experimental results show that the proposed algorithm enables significant transient suppression for a variety of transient types.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.3
2012 A sparse blocking matrix for multiple constraints GSC beamformer
abstract
Modern high performance speech processing applications incorporate large microphone arrays. Complicated scenarios comprising multiple sources, motivate the use of the linearly constrained minimum variance (LCMV) beamformer (BF) and specifically its efficient generalized sidelobe canceler (GSC) implementation. The complexity of applying the GSC is dominated by the blocking matrix (BM). A common approach for constructing the BM is to use a projection matrix to the null-subspace of the constraints. The latter BM is denoted as the eigen-space BM, and requires M2complex multiplications, where M is the number of microphones. In the current contribution, a novel systematic scheme for constructing a multiple constraints sparse BM is presented. The sparsity of the proposed BM substantially reduces the complexity to K × (M - K) complex multiplications, where K is the number of constraints. A theoretical analysis of the signal leakage and of the blocking ability of the proposed sparse BM and of the eigen-space BM is derived. It is proven analytically, and tested for narrowband signals and for speech signals, that the blocking abilities of the sparse and of the eigen-space BMs are equivalent.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
ICASSP2
2012 Supervised Graph-Based Processing for Sequential Transient Interference Suppression
abstract
In this paper, we present a supervised graph-based framework for sequential processing and employ it to the problem of transient interference suppression. Transients typically consist of an initial peak followed by decaying short-duration oscillations. Such sounds, e.g., keyboard typing and door knocking, often arise as an interference in everyday applications: hearing aids, hands-free accessories, mobile phones, and conference-room devices. We describe a graph construction using a noisy speech signal and training recordings of typical transients. The main idea is to capture the transient interference structure, which may emerge from the construction of the graph. The graph parametrization is then viewed as a data-driven model of the transients and utilized to define a filter that extracts the transients from noisy speech measurements. Unlike previous transient interference suppression studies, in this work the graph is constructed in advance from training recordings. Then, the graph is extended to newly acquired measurements, providing a sequential filtering framework of noisy speech.
Ronen Talmon, Israel Cohen, Sharon Gannot, Ronald R. Coifman
IEEE Trans. Speech Audio Process.3
2011 Performance analysis of a randomly spaced wireless microphone array
abstract
A randomly distributed microphone array is considered in this work. In many applications exact design of the array is impractical. The performance of these arrays, characterized by a large number of microphones deployed in vast areas, cannot be analyzed by traditional deterministic methods. We therefore derive a novel statistical model for performance analysis of the MWF beamformer. We consider the scenario of one desired source and one interfering source arriving from the far-field and impinging on a uniformly distributed linear array. A theoretical model for the MMSE is developed and verified by simulations. The applicability of the proposed statistical model for speech signals is discussed.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
ICASSP2
2011 Direction-of-arrival estimation using acoustic vector sensors in the presence of noise
abstract
A vector-sensor consisting of a monopole sensor collocated with orthogonally oriented dipole sensors can be used for direction-of arrival (DOA) estimation. A method is proposed to estimate the DOA based on the direction of maximum power. Algorithms mentioned in earlier works are shown to be special cases of the proposed method. An iterative algorithm based on the principal of gradient ascent is presented for the solution of the maximum power problem. The proposed maximum-power method is shown to approach the Cramer-Rao lower bound (CRLB) with a suitable choice of parameter.
Dovid Levin, Sharon Gannot, Emanuël A. P. Habets
ICASSP2
2011 Clustering and suppression of transient noise in speech signals using diffusion maps
abstract
Recently we have presented a novel approach for transient noise reduction that relies on non-local (NL) filtering. In this paper, we modify and extend our approach to support clustering and suppression of a few transient noise types simultaneously, by introducing two novel concepts. We observe that voiced speech spectral components are slowly varying compared to transient noise. Thus, by applying an algorithm for noise power spectral density (PSD) estimation, configured to track faster variations than pseudo-stationary noise, the PSD of speech components may be estimated. In addition, we utilize diffusion maps to embed the measurements into a new do main. We obtain a new representation which enables clustering of different transient noise types. The new representation is incorporated into a NL filter as a better affinity metric for averaging over transient instances. Experimental results show that the proposed algorithm enables clustering and suppression of multiple transient interferences.
Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP3
2011 Multiple-Hypothesis Extended Particle Filter for Acoustic Source Localization in Reverberant Environments
abstract
Particle filtering has been shown to be an effective approach to solving the problem of acoustic source localization in reverberant environments. In reverberant environment, the direct- arrival of the single source is accompanied by multiple spurious arrivals. Multiple-hypothesis model associated with these arrivals can be used to alleviate the unreliability often attributed to the acoustic source localization problem. Until recently, this multiple- hypothesis approach was only applied to bootstrap-based particle filter schemes. Recently, the extended Kalman particle filter (EPF) scheme which allows for an improved tracking capability was proposed for the localization problem. The EPF scheme utilizes a global extended Kalman filter (EKF) which strongly depends on prior knowledge of the correct hypotheses. Due to this, the extension of the multiple-hypothesis model for this scheme is not trivial. In this paper, the EPF scheme is adapted to the multiple-hypothesis model to track a single acoustic source in reverberant environments. Our work is supported by an extensive experimental study using both simulated data and data recorded in our acoustic lab. Various algorithms and array constellations were evaluated. The results demonstrate the superiority of the proposed algorithm in both tracking and switching scenarios. It is further shown that splitting the array into several sub-arrays improves the robustness of the estimated source location.
A. Levy, Sharon Gannot, Emanuël A. P. Habets
IEEE Trans. Speech Audio Process.2
2011 Transient Noise Reduction Using Nonlocal Diffusion Filters
abstract
Enhancement of speech signals for hands-free communication systems has attracted significant research efforts in the last few decades. Still, many aspects and applications remain open and require further research. One of the important open problems is the single-channel transient noise reduction. In this paper, we present a novel approach for transient noise reduction that relies on non-local (NL) neighborhood filters. In particular, we propose an algorithm for the enhancement of a speech signal contaminated by repeating transient noise events. We assume that the time duration of each reoccurring transient event is relatively short compared to speech phonemes and model the speech source as an auto-regressive (AR) process. The proposed algorithm consists of two stages. In the first stage, we estimate the power spectral density (PSD) of the transient noise by employing a NL neighborhood filter. In the second stage, we utilize the optimally modified log spectral amplitude (OM-LSA) estimator for denoising the speech using the noise PSD estimate from the first stage. Based on a statistical model for the measurements and diffusion interpretation of NL filtering, we obtain further insight into the algorithm behavior. In particular, for given transient noise, we determine whether estimation of the noise PSD is feasible using our approach, how to properly set the algorithm parameters, and what is the expected performance of the algorithm. Experimental study shows good results in enhancing speech signals contaminated by transient noise, such as typical household noises, construction sounds, keyboard typing, and metronome clacks.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.3
2010 Subspace tracking of multiple sources and its application to speakers extraction
abstract
In this paper we introduce a novel algorithm for extracting desired speech signals uttered by moving speakers contaminated by competing speakers and stationary noise in a reverberant environment. The proposed beamformer uses eigenvectors spanning the desired and interference signals subspaces. It relaxes the common requirement on the activity patterns of the various sources. A novel mechanism for tracking the desired and interferences subspaces is proposed, based on the projection approximation subspace tracking (deflation) (PASTd) procedure and on a union of subspaces procedure. This contribution extends previously proposed methods to deal with multiple speakers in dynamic scenarios.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
ICASSP2
2010 Speech enhancement in transient noise environment using diffusion filtering
abstract
Recently, we have presented a transient noise reduction algorithm for speech signals that relies on non-local diffusion filtering. By exploiting the repetitive nature of transient noises we proposed a simple and efficient algorithm, which enabled suppression of various noise types. In this paper, we incorporate a modified diffusion operator in order to obtain a more robust algorithm and further enhancement of the speech. We demonstrate the performance of the modified algorithm and compare it with a competing solution. We show that the proposed algorithm enables improved suppression of various transient interferences without any further computational burden.
Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP3
2010 New Insights Into the MVDR Beamformer in Room Acoustics
abstract
The minimum variance distortionless response (MVDR) beamformer, also known as Capon's beamformer, is widely studied in the area of speech enhancement. The MVDR beamformer can be used for both speech dereverberation and noise reduction. This paper provides new insights into the MVDR beamformer. Specifically, the local and global behavior of the MVDR beamformer is analyzed and novel forms of the MVDR filter are derived and discussed. In earlier works it was observed that there is a tradeoff between the amount of speech dereverberation and noise reduction when the MVDR beamformer is used. Here, the tradeoff between speech dereverberation and noise reduction is analyzed thoroughly. The local and global behavior, as well as the tradeoff, is analyzed for different noise fields such as, for example, a mixture of coherent and non-coherent noise fields, entirely non-coherent noise fields and diffuse noise fields. It is shown that maximum noise reduction is achieved when the MVDR beamformer is used for noise reduction only. The amount of noise reduction that is sacrificed when complete dereverberation is required depends on the direct-to-reverberation ratio of the acoustic impulse response between the source and the reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction. When desiring both speech dereverberation and noise reduction, the results also demonstrate that the amount of noise reduction that is sacrificed decreases when the number of microphones increases.
Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot, Jacek Dmochowski
IEEE Trans. Speech Audio Process.4
2009 On a tradeoff between dereverberation and noise reduction using the MVDR beamformer
abstract
The minimum variance distortionless response (MVDR) beamformer can be used for both speech dereverberation and noise reduction. In this paper we analyse the tradeoff between the amount of speech dereverberation and noise reduction achieved by the MVDR beamformer. We show that the amount of noise reduction that is sacrificed when desiring both speech dereverberation and noise reduction depends on the direct-to-reverberation ratio of the acoustic transfer function between the desired source and a reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction.
Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot
ICASSP4
2009 Multichannel speech enhancement using convolutive transfer function approximation in reverberant environments
abstract
Recently, we have presented a transfer-function generalized sidelobe canceler (TF-GSC) beamformer in the short time Fourier transform domain, which relies on a convolutive transfer function approximation of relative transfer functions between distinct sensors. In this paper, we combine a delay-and-sum beamformer with the TF-GSC structure in order to suppress the speech signal reflections captured at the sensors in reverberant environments. We demonstrate the performance of the proposed beamformer and compare it with the TF-GSC. We show that the proposed algorithm enables suppression of reverberations and further noise reduction compared with the TF-GSC beamformer.
Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP3
2009 Late Reverberant Spectral Variance Estimation Based on a Statistical Model
abstract
In speech communication systems the received microphone signals are degraded by room reverberation and ambient noise that decrease the fidelity and intelligibility of the desired speaker. Reverberant speech can be separated into two components, viz. early speech and late reverberant speech. Recently, various algorithms have been developed to suppress late reverberant speech. One of the main challenges is to develop an estimator for the so-called late reverberant spectral variance (LRSV) which is required by most of these algorithms. In this letter a statistical reverberation model is proposed that takes the energy contribution of the direct-path into account. This model is then used to derive a more general LRSV estimator, which in a particular case reduces to an existing LRSV estimator. Experimental results show that the developed estimator is advantageous in case the source-microphone distance is smaller than the critical distance.
Emanuël A. P. Habets, Sharon Gannot, Israel Cohen
IEEE Signal Process. Lett.2
2009 Multichannel Eigenspace Beamforming in a Reverberant Noisy Environment With Multiple Interfering Speech Signals
abstract
In many practical environments we wish to extract several desired speech signals, which are contaminated by nonstationary and stationary interfering signals. The desired signals may also be subject to distortion imposed by the acoustic room impulse responses (RIRs). In this paper, a linearly constrained minimum variance (LCMV) beamformer is designed for extracting the desired signals from multimicrophone measurements. The beamformer satisfies two sets of linear constraints. One set is dedicated to maintaining the desired signals, while the other set is chosen to mitigate both the stationary and nonstationary interferences. Unlike classical beamformers, which approximate the RIRs as delay-only filters, we take into account the entire RIR [or its respective acoustic transfer function (ATF)]. The LCMV beamformer is then reformulated in a generalized sidelobe canceler (GSC) structure, consisting of a fixed beamformer (FBF), blocking matrix (BM), and adaptive noise canceler (ANC). It is shown that for spatially white noise field, the beamformer reduces to a FBF, satisfying the constraint sets, without power minimization. It is shown that the application of the adaptive ANC contributes to interference reduction, but only when the constraint sets are not completely satisfied. We show that relative transfer functions (RTFs), which relate the desired speech sources and the microphones, and a basis for the interference subspace suffice for constructing the beamformer. The RTFs are estimated by applying the generalized eigenvalue decomposition (GEVD) procedure to the power spectral density (PSD) matrices of the received signals and the stationary noise. A basis for the interference subspace is estimated by collecting eigenvectors, calculated in segments where nonstationary interfering sources are active and the desired sources are inactive. The rank of the basis is then reduced by the application of the orthogonal triangular decomposition (QRD). This procedure relaxes the common requirement for nonoverlapping activity periods of the interference sources. A comprehensive experimental study in both simulated and real environments demonstrates the performance of the proposed beamformer.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.2
2009 Relative Transfer Function Identification Using Convolutive Transfer Function Approximation
abstract
In this paper, we present a relative transfer function (RTF) identification method for speech sources in reverberant environments. The proposed method is based on the convolutive transfer function (CTF) approximation, which enables to represent a linear convolution in the time domain as a linear convolution in the short-time Fourier transform (STFT) domain. Unlike the restrictive and commonly used multiplicative transfer function (MTF) approximation, which becomes more accurate when the length of a time frame increases relative to the length of the impulse response, the CTF approximation enables representation of long impulse responses using short time frames. We develop an unbiased RTF estimator that exploits the nonstationarity and presence probability of the speech signal and derive an analytic expression for the estimator variance. Experimental results show that the proposed method is advantageous compared to common RTF identification methods in various acoustic environments, especially when identifying long RTFs typical to real rooms.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.3
2009 Convolutive Transfer Function Generalized Sidelobe Canceler
abstract
In this paper, we propose a convolutive transfer function generalized sidelobe canceler (CTF-GSC), which is an adaptive beamformer designed for multichannel speech enhancement in reverberant environments. Using a complete system representation in the short-time Fourier transform (STFT) domain, we formulate a constrained minimization problem of total output noise power subject to the constraint that the signal component of the output is the desired signal, up to some prespecified filter. Then, we employ the general sidelobe canceler (GSC) structure to transform the problem into an equivalent unconstrained form by decoupling the constraint and the minimization. The CTF-GSC is obtained by applying a convolutive transfer function (CTF) approximation on the GSC scheme, which is a more accurate and a less restrictive than a multiplicative transfer function (MTF) approximation. Experimental results demonstrate that the proposed beamformer outperforms the transfer function GSC (TF-GSC) in reverberant environments and achieves both improved noise reduction and reduced speech distortion.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.3
2008 Dual-microphone speech dereverberation using GARCH modeling
abstract
In this paper, we develop a dual-microphone speech dereverberation algorithm for noisy environments, which is aimed at suppressing late reverberation and background noise. The spectral variance of the late reverberation is obtained with adaptively-estimated direct path compensation. A Markov-switching generalized autoregressive conditional heteroscedasticity (GARCH) model is used to estimate the spectral variance of the desired signal, which includes the direct sound and early reverberation. Experimental results demonstrate the advantage of the proposed algorithm compared to a decision-directed-based algorithm.
Ari Abramson, Emanuël A. P. Habets, Sharon Gannot, Israel Cohen
ICASSP3
2008 Performance bounds for channel tracking algorithms For MIMO systems
abstract
In this paper we derive performance bounds for tracking time-varying OFDM multiple-input multiple-output (MIMO) communication channel in the presence of additive white Gaussian noise (AWGN). We discuss two channel tracking schemes. The first tracks the filter coefficients directly in time-domain, while the second separately tracks each tone in the frequency-domain. The Kalman filter, with known channel statistics, is utilized for evaluating the performance bounds. It is shown that the time-domain tracking scheme, which exploits the sparseness of the channel impulse response, outperforms the computationally more efficient, frequency-domain tracking scheme, which does not exploit the smooth frequency response of the channel.
Livnat Ehrenberg, Sharon Gannot, Amir Leshem, Ephraim Zehavi
ICASSP2
2008 Joint Dereverberation and Residual Echo Suppression of Speech Signals in Noisy Environments
abstract
Hands-free devices are often used in a noisy and reverberant environment. Therefore, the received microphone signal does not only contain the desired near-end speech signal but also interferences such as room reverberation that is caused by the near-end source, background noise and a far-end echo signal that results from the acoustic coupling between the loudspeaker and the microphone. These interferences degrade the fidelity and intelligibility of near-end speech. In the last two decades, postfilters have been developed that can be used in conjunction with a single microphone acoustic echo canceller to enhance the near-end speech. In previous works, spectral enhancement techniques have been used to suppress residual echo and background noise for single microphone acoustic echo cancellers. However, dereverberation of the near-end speech was not addressed in this context. Recently, practically feasible spectral enhancement techniques to suppress reverberation have emerged. In this paper, we derive a novel spectral variance estimator for the late reverberation of the near-end speech. Residual echo will be present at the output of the acoustic echo canceller when the acoustic echo path cannot be completely modeled by the adaptive filter. A spectral variance estimator for the so-called late residual echo that results from the deficient length of the adaptive filter is derived. Both estimators are based on a statistical reverberation model. The model parameters depend on the reverberation time of the room, which can be obtained using the estimated acoustic echo path. A novel postfilter is developed which suppresses late reverberation of the near-end speech, residual echo and background noise, and maintains a constant residual background noise level. Experimental results demonstrate the beneficial use of the developed system for reducing reverberation, residual echo, and background noise.
Emanuël A. P. Habets, Sharon Gannot, Israel Cohen, P. Sommen
IEEE Trans. Speech Audio Process.2
2008 Dual-Source Transfer-Function Generalized Sidelobe Canceller
abstract
Full-duplex hands-free man/machine interface often suffers from directional nonstationary interference, such as a competing speaker, as well as stationary interferences which may comprise both directional and nondirectional signals. The transfer-function generalized sidelobe canceller (TF-GSC) exploits the nonstationarity of the speech signal to enhance it when the undesired interfering signals are stationary. Unfortunately, the assumptions leading to the derivation of the TF-GSC are violated when a nonstationary interference is present. In this paper, we propose an adaptive beamformer, based on the TF-GSC, that is suitable for cancelling nonstationary interferences in noisy reverberant environments. We modify two of the TF-GSC components to enable suppression of the nonstationary undesired signal. A modified fixed beamformer (FBF) is designed to block the nonstationary interfering signal while maintaining the desired speech signal. A modified blocking matrix (BM) is designed to block both the desired signal and the nonstationary interference. We introduce a novel method for updating the blocking matrix in double talk scenarios, which exploits the nonstationarity of both the desired and interfering speech signals. Experimental results demonstrate the performance of the proposed algorithm in noisy and reverberant environments and show its superiority over the original TF-GSC.
Gal Reuven, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.2
2007 Dual-Microphone Speech Dereverberation using a Reference Signal
abstract
Speech signals recorded with a distant microphone usually contain reverberation, which degrades the fidelity and intelligibility of speech, and the recognition performance of automatic speech recognition systems. In this paper we propose a speech dereverberation system which uses two microphones. A generalized sidelobe canceller (GSC) type of structure is used to enhance the desired speech signal. The GSC structure is used to create two signals. The first signal is the output of a standard delay and sum beamformer, and the second signal is a reference signal which is constructed such that the direct speech signal is blocked. We propose to utilize the reverberation which is present in the reference signal to enhance the output of the delay and sum beamformer. The power envelope of the reference signal and the power envelope of the output of the delay and sum beamformer are used to estimate the residual reverberation in the output of the delay and sum beamformer. The output of the delay and sum beamformer is then enhanced using a spectral enhancement technique. The proposed method only requires an estimate of the direction of arrival of the desired speech source. Experiments using simulated room impulse responses are presented and show significant reverberation reduction while keeping the speech distortion low.
Emanuël A. P. Habets, Sharon Gannot
ICASSP (4)2
2007 Multichannel Acoustic Echo Cancellation and Noise Reduction in Reverberant Environments using the Transfer-Function GSC
abstract
In this paper, we present a multi-channel acoustic echo canceller that is integrated into the transfer-function generalized sidelobe canceller (TF-GSC). The proposed scheme consists of a primary TF-GSC, for dealing with the noise interferences, and a secondary modified TF-GSC, for dealing with the echo cancellation. The secondary TF-GSC includes an echo canceller embedded within a replica of the primary TF-GSC components. Experimental results demonstrate improved performance compared to cascade schemes of acoustic echo cancellation and adaptive beamforming.
Gal Reuven, Sharon Gannot, Israel Cohen
ICASSP (1)2
2007 Special issue on Speech Enhancement
Philipos C. Loizou, Israel Cohen, Sharon Gannot, Kuldip K. Paliwal
Speech Commun.3
2007 Performance analysis of dual source transfer-function generalized sidelobe canceller
Gal Reuven, Sharon Gannot, Israel Cohen
Speech Commun.2
2007 Joint noise reduction and acoustic echo cancellation using the transfer-function generalized sidelobe canceller
Gal Reuven, Sharon Gannot, Israel Cohen
Speech Commun.2
2005 Time difference of arrival estimation of speech source in a noisy and reverberant environment
Tsvi G. Dvorkind, Sharon Gannot
Signal Process.2
2004 Speech enhancement based on the general transfer function GSC and postfiltering
abstract
In speech enhancement applications microphone array postfiltering allows additional reduction of noise components at a beamformer output. Among microphone array structures the recently proposed general transfer function generalized sidelobe canceller (TF-GSC) has shown impressive noise reduction abilities in a directional noise field, while still maintaining low speech distortion. However, in a diffused noise field less significant noise reduction is obtainable. The performance is even further degraded when the noise signal is nonstationary. In this contribution we propose three postfiltering methods for improving the performance of microphone arrays. Two of which are based on single-channel speech enhancers and making use of recently proposed algorithms concatenated to the beamformer output. The third is a multichannel speech enhancer which exploits noise-only components constructed within the TF-GSC structure. This work concentrates on the assessment of the proposed postfiltering structures. An extensive experimental study, which consists of both objective and subjective evaluation in various noise fields, demonstrates the advantage of the multichannel postfiltering compared to the single-channel techniques.
Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.1
2003 Speech enhancement based on the general transfer function GSC and postfiltering
abstract
In speech enhancement applications, microphone array postfiltering allows additional reduction of noise components at a beamformer output. Among microphone array structures, the recently proposed general transfer function generalized sidelobe canceller (TF-GSC) has shown impressive noise reduction abilities in a directional noise field, while still maintaining low speech distortion. However, in a diffused noise field, less significant noise reduction is obtainable. The performance is even further degraded when the noise is nonstationary. We present three postfiltering methods for improving the performance of microphone arrays. Two of them are based on single-channel speech enhancers and make use of recently proposed algorithms concatenated to the beamformer output. The third is a multichannel speech enhancer which exploits noise-only components constructed within the TF-GSC structure. An experimental study, which consists of both objective and subjective evaluation in various noise fields, demonstrates the advantage of the multi-channel postfiltering compared to single-channel techniques.
Sharon Gannot, Israel Cohen
ICASSP (1)1
2002 Speech enhancement using a mixture-maximum model
abstract
We present a spectral domain, speech enhancement algorithm. The new algorithm is based on a mixture model for the short time spectrum of the clean speech signal, and on a maximum assumption in the production of the noisy speech spectrum. In the past this model was used in the context of noise robust speech recognition. In this paper we show that this model is also effective for improving the quality of speech signals corrupted by additive noise. The computational requirements of the algorithm can be significantly reduced, essentially without paying performance penalties, by incorporating a dual codebook scheme with tied variances. Experiments, using recorded speech signals and actual noise sources, show that in spite of its low computational requirements, the algorithm shows improved performance compared to alternative speech enhancement algorithms.
David Burshtein, Sharon Gannot
IEEE Trans. Speech Audio Process.2
1999 Speech enhancement using a mixture-maximum model
David Burshtein, Sharon Gannot
EUROSPEECH2
1998 Iterative and sequential Kalman filter-based speech enhancement algorithms
abstract
Speech quality and intelligibility might significantly deteriorate in the presence of background noise, especially when the speech signal is subject to subsequent processing. In particular, speech coders and automatic speech recognition (ASR) systems that were designed or trained to act on clean speech signals might be rendered useless in the presence of background noise. Speech enhancement algorithms have therefore attracted a great deal of interest. In this paper, we present a class of Kalman filter-based algorithms with some extensions, modifications, and improvements of previous work. The first algorithm employs the estimate-maximize (EM) method to iteratively estimate the spectral parameters of the speech and noise parameters. The enhanced speech signal is obtained as a byproduct of the parameter estimation algorithm. The second algorithm is a sequential, computationally efficient, gradient descent algorithm. We discuss various topics concerning the practical implementation of these algorithms. Extensive experimental study using real speech and noise signals is provided to compare these algorithms with alternative speech enhancement algorithms, and to compare the performance of the iterative and sequential algorithms.
Sharon Gannot, David Burshtein, Ehud Weinstein
IEEE Trans. Speech Audio Process.1
1997 Iterative-batch and sequential algorithms for single microphone speech enhancement
abstract
Speech quality and intelligibility might significantly deteriorate in the presence of background noise, especially when the speech signal is subject to subsequent processing. In this paper we represent a class of Kalman-filter based speech enhancement algorithms with some extensions, modifications, and improvements. The first algorithm employs the estimate-maximize (EM) method to iteratively estimate the spectral parameters of the speech and noise parameters. The enhanced speech signal is obtained as a by-product of the parameter estimation algorithm. The second algorithm is a sequential, computationally efficient, gradient descent algorithm. We discuss various topics concerning the practical implementation of these algorithms. Experimental study, using real speech and noise signals is provided to compare these algorithms with alternative speech enhancement algorithms, and to compare the performance of the iterative and sequential algorithms.
Sharon Gannot, David Burshtein, Ehud Weinstein
ICASSP1