Israel Cohen

dblp:61/2149 · DBLP profile ↗
← Back
159ranked-venue papers
27as first author
40since 2021 · last 2025
0000-0002-2556-3972ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 106 · 22 first-author · 26 since 2021Artificial intelligence and machine learning · 54 · 8 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 E-URES 2.0: Efficient User-Centric Residual-Echo Suppression with a Lightweight Neural Network
abstract
We recently introduced the Efficient User-centric Residual-Echo Suppression (E-URES) framework, which significantly reduces the floating-point operations per second (FLOPS) required during inference by 90% compared to the URES framework. The E-URES operates based on a user-operating point (UOP) defined by two key metrics: the residual echo suppression level (RESL) and the desired-speech maintained level (DSML) that the user anticipates from the output signal of a residual echo suppression (RES) system. In the first stage, an ensemble of 101 branches is employed, where each branch has two cascaded neural networks: a preliminary RES system with a design parameter, which varies between branches and balances the RESL and DSML of its RES systems’ prediction, and a subsequent UOP estimator. In the second stage, a neural network uses available acoustic signals and the UOP to predict which three branches achieve the highest acoustic echo cancellation mean opinion score (AECMOS) within a specified UOP-error tolerance. Then, costly AECMOS calculations are performed only for these selected branches. Despite this efficiency mechanism, the E-URES can apply real-time inference only with dedicated and expensive hardware, limiting its wide adoption. Here, we present E-URES 2.0, which focuses on reducing the computational costs of E-URES in its first stage. A lightweight neural network preprocesses available acoustic signals and the UOP to track a subset of the 101 design parameters that their branches produce the most accurate UOP estimations in their outcomes. Only these branches are calculated during inference and continue to the AECMOS estimation stage. With 60 hours of data, we show that with a negligible performance drop on average, the E-URES 2.0 can reduce 87% of the branches and 61% of the FLOPS of the E-URES and can achieve real-time inference with standard, affordable hardware.
Amir Ivry, Israel Cohen
ICASSP2
2025 A Bilinear Source Separation, Dereverberation, and Background Noise Suppression Algorithm for Augmented Reality Applications
abstract
This paper introduces a new bilinear algorithm that combines distortionless response beamforming with weighted prediction error using the Kronecker product operator to improve speech signal quality in different acoustic environments. The algorithm utilizes recursive least squares cost functions and handles real dynamic, noisy, and reverberant conditions effectively. Compared to a recently proposed approach, our approach outperforms separating the desired source from undesired sources while providing dereverberation and noise suppression. It is also preferable regarding the quality of the enhanced signals it produces. In contrast, it exhibits a more considerable desired signal distortion and reduced intelligibility. The algorithm’s computational efficiency and robustness make it suitable for real-time applications, as validated using real recordings from the SPeech Enhancement for Augmented Reality (SPEAR) challenge.
Alon Nemirovsky, Gal Itzhak, Israel Cohen
ICASSP3
2025 MBSS-T1: Model-based subject-specific self-supervised motion correction for robust cardiac T1 mapping
Eyal Hanania, Adi Zehavi-Lenz, Ilya Volovik, Daphna Link-Sourani, Israel Cohen, Moti Freiman
Medical Image Anal.5
2024 Array Geometry Optimization for Region-of-Interest Near-Field Beamforming
abstract
Microphone array geometry plays a crucial role in array signal processing and beamforming. This paper presents an array geometry optimization approach for near-field beamforming. Considering a continuous region of interest, we aim to maximize the worst-case broadband directivity of the array while ensuring sufficiently high white noise gain. We use a near-field wave propagation analysis to formulate a convex optimization problem and find the optimal linear array topology. We evaluate the performance of the proposed approach for different desired signal directions and distances and compare our results to traditional beamformers. The optimized non-uniform geometry achieves better white noise gain and higher array directivity than standard uniform linear arrays. Furthermore, the proposed method can be used for near- and far-field sources.
Ron Moisseev, Gal Itzhak, Israel Cohen
ICASSP3
2024 Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
Yehoshua Dissen, Shiry Yonash, Israel Cohen, Joseph Keshet
INTERSPEECH3
2024 Learning Noise Adapters for Incremental Speech Enhancement
abstract
Incremental speech enhancement (ISE), with the ability to incrementally adapt to new noise domains, represents a critical yet comparatively under-investigated topic. While the regularization-based method has been proposed to solve the ISE task, it usually suffers from the dilemma wherein the gain of one domain directly entails the loss of another. To solve this issue, we propose an effective paradigm, termed Learning Noise Adapters (LNA), which significantly mitigates the catastrophic domain forgetting phenomenon in the ISE task. In our methodology, we employ a frozen pre-trained model to train and retain a domain-specific adapter for each newly encountered domain, enabling the capture of variations in feature distributions within these domains. Subsequently, our approach involves the development of an unsupervised, training-free noise selector for the inference stage, which is responsible for identifying the domains of test speech samples. A comprehensive experimental validation has substantiated the effectiveness of our approach.
Ziye Yang, Xiang Song 0005, Jie Chen 0022, Cédric Richard, Israel Cohen
IEEE Signal Process. Lett.5
2024 Howling Detection and Gain Control for Speech Reinforcement in a Noisy Car Cabin Environment
abstract
In-car speech communication is particularly challenging due to environmental noise. The speaker's microphone also acquires car and road noises, resulting in a low signal-to-noise ratio and persistent frequency-howls that do not decrease, which degrade the system's output sound quality. In this paper, we address the problem of howling control for in-room speech reinforcement systems under high environmental noise levels, specifically in a car cabin. We control the acoustic feedback via a gain control algorithm based on howling detection. Two detection methods are proposed to detect non-decreasing underdamped frequency-howls. A single-channel optimal Feedback Wiener gain is derived to enhance the desired near-end speech signal, followed by another Wiener filter that is based on the signal magnitude relative to the estimated noise power spectral density. Adjusting the howling energy threshold is proposed to deal with false-detection artifacts arising as the environment's environmental noise level rises. The performance improvements of the howling-detection-based gain control algorithm following the proposed adjustments are evaluated in clean and noisy environments. As detection of non-decreasing underdamped frequency-howls is feasible, it was found that it becomes unnecessary when the environmental noise is successfully reduced via the Wiener filter.
Yehav Alkaher, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 A User-Centric Approach for Deep Residual-Echo Suppression in Double-Talk
abstract
We introduce a user-centric residual-echo suppression (URES) framework in double-talk. This framework receives a user operating point (UOP) that consists of two metric values: the residual echo suppression level (RESL) and the desired speech-maintained level (DSML) that the user expects from the RES outcome. Then, the URES pipeline undergoes three stages. Firstly, we consider a deep RES model with a tunable design parameter that balances between the RESL and DSML and utilizes 101 pre-trained instances of this model, each with a different design parameter value. Thus, an identical input is expected to generate a different pair of RESL and DSML values in the prediction of every instance. Second, every prediction is separately fed to a subsequent pre-trained deep model instance that estimates the RESL and DSML of the prediction since these metrics depend on unavailable information in practice. Lastly, each pair of RESL and DSML estimates is compared with the UOP. The pairs that match the UOP up to a given tolerance threshold are narrowed down to the prediction with the maximal acoustic-echo cancellation mean-opinion score (AECMOS), which is the output of the URES system. This suggested framework holds three prominent advantages introduced in this study: it generates an RES output with RESL and DSML that match a UOP, supports near-real-time tracking of UOP changes, and applies AECMOS maximization. Experimental results consider 60 h of varied real and synthetic data. Average results can achieve an AECMOS subjectively considered excellent with RESL and DSML deviations of roughly 2 dB from the UOP. Any UOP adjustment can be tracked in less than 40 ms with a real-time factor of 1.92, but due to the high computational resources demanded by the framework, this is enabled on-edge only with high-end dedicated hardware, which limits general availability.
Amir Ivry, Israel Cohen, Baruch Berdugo
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Constant Elevation-Beamwidth Beamforming With Concentric Ring Arrays
abstract
A hybrid approach is proposed to efficiently design a constant elevation-beamwidth beamforming with concentric ring arrays (CRAs). The design exploits the degrees of freedom of the array geometry for superior performance. In particular, the ring radii and the beamformer coefficients are optimized simultaneously for all frequencies. We introduce a convex quadratic programming problem for a given CRA configuration, directly optimizing the beamformer coefficients while maintaining constant elevation-beamwidth over a wide range of frequencies. The proposed objective function contains a control variable, allowing a tradeoff between the directivity factor and the white noise gain. Subsequently, a hybrid approach is proposed to optimize the ring radii with a genetic algorithm exploiting the partial convexity of the problem. Experimental results demonstrate the flexibility and advantages of the proposed approach compared to the state-of-the-art in terms of directivity factor, white noise gain, sidelobe level, and beamwidth consistency, with reduced resources and a significantly lower computation time.
Orel Peretz, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech Dereverberation
abstract
Dereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those methods generally assume that there is only a single speaker in the acoustic environment and, consequently, they suffer from significant performance degradation if multiple speakers participate in the conversation. How to deal with reverberation in multiple-speaker scenarios is still a challenging problem, which is studied in this work. We present a switching multichannel linear prediction filtering method, which designs multiple linear filters with each tracking one speaker. When some speaker is active, the corresponding filter and the weighted cross-correlation matrix are updated while the other filters are kept unchanged. To further improve the performance and reduce complexity, we apply the Kronecker product to decompose every linear prediction filter into a Kronecker product of two shorter filters: one is time-invariant and the other is time-varying. The former is estimated with a batch method (using only a few seconds of speech signal when the corresponding speaker starts to talk in the entire conversation) while a recursive least-squares algorithm is derived for identifying the time-varying set of Kronecker filters.
Gongping Huang, Jacob Benesty, Israel Cohen, Emil Winebrand, Jingdong Chen, Walter Kellermann
ICASSP3
2023 PCMC-T1: Free-Breathing Myocardial T1 Mapping with Physically-Constrained Motion Correction
Eyal Hanania, Ilya Volovik, Lilach Barkat, Israel Cohen, Moti Freiman
MICCAI (7)4
2023 Differentiable Mean Opinion Score Regularization for Perceptual Speech Enhancement
Tomer Rosenbaum, Israel Cohen, Emil Winebrand, Ofri Gabso
Pattern Recognit. Lett.2
2023 Differential constant-beamwidth beamforming with cube arrays
Gal Itzhak, Israel Cohen
Speech Commun.2
2022 Incoherent Synthesis of Sparse Broadband Arrays based on a Parameter-Free Subspace Clustering
abstract
In this paper, we propose an incoherent design method of sparse broadband arrays that optimizes the number of sensors and their positions simultaneously. We introduce an iterative clustering procedure that merges different groups of sensors with a small distance, in terms of Bhattacharyya distance, between their angle distributions. The iterative clustering procedure is initialized with a large number of groups of sensors, and computes in each iteration a clustering score and a threshold. Then, near groups are merged into joint groups, yielding a new set of groups of sensors. We show that the optimal set of sensors is obtained when the clustering score is larger than the threshold, indicating that the remaining groups are distant. The proposed approach is demonstrated by a design of a superdirective beamformer, and its performance is compared with an existing incoherent approach. Experimental results show improved performance in terms of a more favorable tradeoff between directivity factor and white noise gain.
Guy Gubnitsky, Yaakov Buchris, Israel Cohen
ICASSP3
2022 Deep Adaptation Control for Acoustic Echo Cancellation
abstract
We propose a general framework for adaptation control using deep neural networks (NNs) and apply it to acoustic echo cancellation (AEC). First, the optimal step-size that controls the adaptation is derived offline by solving a constrained nonlinear optimization problem that minimizes the adaptive filter misadjustment. Then, a deep NN is trained to learn the relation between the input data and the optimal step-size. In real-time, the NN infers the optimal step-size from streaming data and feeds it to an NLMS filter for AEC. This data-driven method makes no assumptions on the acoustic setup and is entirely non-parametric. Experiments with 100 h of real and synthetic data show that the proposed method outperforms the competition in echo cancellation, speech distortion, and convergence during both single-talk and double-talk.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP2
2022 Off-the-Shelf Deep Integration For Residual-Echo Suppression
abstract
Residual-echo suppression (RES) systems suppress the echo and preserve the speech from a mixture of the two. In hands-free speech communication, RES may also be addressed as a source separation (SS) or speech enhancement (SE) problem, where the echo can be manipulated as an interfering speech signal. In this study, we fine-tune three pre-trained deep learning-based systems originally designed for RES, SS, and SE, and show that the best performing system for the task of RES varies with respect to the acoustic conditions. Then, we propose a real-time data-driven integration of these systems, where a neural network continuously tracks the system that achieves the best performance during both single-talk and double-talk periods. Experiments with 100 h of real and synthetic data show that the integrated system outperforms each individual system in terms of echo suppression and speech distortion in various acoustic environments.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP2
2022 Attenuation Of Acoustic Early Reflections In Television Studios Using Pretrained Speech Synthesis Neural Network
abstract
Machine learning and digital signal processing have been extensively used to enhance speech. However, methods to reduce early reflections in studio settings are usually related to the physical characteristics of the room. In this paper, we address the problem of early acoustic reflections in television studios and control rooms, and propose a two-stage method that exploits the knowledge of a pretrained speech synthesis generator. First, given a degraded speech signal that includes the direct sound and early reflections, a U-Net convolutional neural network is used to attenuate the early reflections in the spectral domain. Then, a pretrained speech synthesis generator reconstructs the phase to predict an enhanced speech signal in the time domain. Qualitative and quantitative experimental results demonstrate excellent studio quality of speech enhancement.
Tomer Rosenbaum, Israel Cohen, Emil Winebrand
ICASSP2
2022 Study of the Null Directions on The Performance of Differential Beamformers
abstract
Null directions are important parameters for differential beamformers, which play an important role on the beamforming performance. In this paper, we investigate the performance of differential beamformers as a function of the null directions. We first derive the directivity factor (DF) as an explicit function of null and show that the DF decreases to 0 if any null approaches to the desired look direction. We then validate the theoretical analysis through simulations using the beampattern, DF and signal-to-interference gain as the performance measures. The results show that: 1) the performance of a differential beamformer degrades significantly if there is any null close to the desired look direction; 2) with a fixed null direction, increasing the order of the differential beamformer can help improve performance.
Xuehan Wang, Israel Cohen, Jacob Benesty, Jingdong Chen
ICASSP2
2022 Intelligent Reflecting Surface OFDM Communication with Deep Neural Prior
abstract
An Intelligent Reflecting Surface (IRS) is an emerging technology for improving the data rate over wireless channels by controlling the underlying channel. In this paper, we describe a novel solution for IRS configuration to maximize the data rate over wideband channels. The optimization is obtained by online training of a deep generative neural network. Inspired by related works in image processing, this network is randomly initialized and acts as a regularization term for the optimization process since the structure of the generator is sufficient to capture a great deal of IRS statistics prior to any learning. In contrast to recent deep learning techniques for IRS configuration, the proposed technique does not require an offline training stage and can adapt quickly to any environment. Compared to the previous state-of-the-art algorithm, the proposed method is significantly faster and obtains IRS configurations that achieve higher data transmission rates.
Tomer Fireaizen, Gal Metzer, Dan Ben-David, Yair Moshe, Israel Cohen, Emil Björnson
ICC5
2022 Challenges and Opportunities in Multi-device Speech Processing
abstract
We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidevice home environments to provide context for the INTER-SPEECH 2022 special session, "Challenges and opportunities for signal processing and machine learning for multiple smart devices".We also identify the datasets needed to support these research areas.Based on the review and our research experience in the multi-device domain, we conclude with an outlook on the future evolution of multiple device signal processing and machine learning.
Gregory Ciccarelli, Jarred Barber, Arun Nair, Israel Cohen
INTERSPEECH4
2022 Objective Metrics to Evaluate Residual-Echo Suppression During Double-Talk in the Stereophonic Case
Amir Ivry, Israel Cohen, Baruch Berdugo
INTERSPEECH2
2022 Single-Sensor Localization of Moving Sources Using Diffusion Kernels and Brownian Motion Model
abstract
Recently, we introduced a single-sensor method for estimating the location and velocity of moving sources. The obtained state-of-the-art results are valid to slow sources whose velocities are gradually changing through time. In this paper, we challenge our algorithm's fundamental assumption. We apply the algorithm to sources that have rapid and random fluctuations in their velocity. The proposed algorithm is based on a supervised learning approach, using diffusion maps with a Euclidean distance-based diffusion kernel. Experimental results demonstrate the benefits of the proposed single-sensor localization method for sources with a Brownian motion model of randomly fluctuating movements.
Eran Zeitouni, Israel Cohen
MMSP2
2022 Time-varying carrier frequency offset estimation in OFDM underwater acoustic communication
Gilad Avrashi, Alon Amar, Israel Cohen
Signal Process.3
2022 Multistage approach for steerable differential beamforming with rectangular arrays
Gal Itzhak, Jacob Benesty, Israel Cohen
Speech Commun.3
2022 Kronecker Product Multichannel Linear Filtering for Adaptive Weighted Prediction Error-Based Speech Dereverberation
abstract
Reverberation, whichis caused by late reflections, impairs not only speech quality but also intelligibility. Consequently, dereverberation, a process to mitigate the impact of reverberation, has attracted significant research interests. Numerous approaches have been developed in the literature, among which the weighted-prediction-error (WPE) one has demonstrated promising potential for reducing or eliminating reverberation. The WPE method has been well studied and several variants have been developed. The adaptive one, called adaptive WPE (AWPE) method, has been widely investigated for use in real applications as it can deal with reverberation in time-varying acoustic environments. However, the computational complexity of AWPE is high, which may be a problem for its implementation in real-time systems. This paper presents some new insights into AWPE-based speech dereverberation by introducing the concepts of Kronecker product and partially time-varying filtering. It then develops two algorithms for dereverberation with lower complexity than AWPE. The significant contributions of this work are as follows. First, we propose a Kronecker product filtering framework for speech dereverberation, where the linear prediction filter is formulated as the Kronecker product of two sets of shorter filters. Second, we propose a partially time-varying Kronecker product filter for dereverberation. Instead of estimating the entire linear prediction filter as in the conventional method, the proposed one only needs to update part of the filter. The proposed approaches can significantly reduce the computational complexity without sacrificing dereverberation performance as compared to AWPE. Simulation results validate the theoretical analysis and justify the advantages of the new methods.
Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Constant-Beamwidth Beamforming With Nonuniform Concentric Ring Arrays
abstract
Solutions for frequency-invariant beamforming with concentric circular arrays are used in many applications, including microphone arrays and audio communication. However, existing methods focus on the azimuth-beamwidth and consider uniformly spaced circular arrays. This paper presents an approach to designing a constant elevation-beamwidth beamformer for nonuniform concentric ring arrays. The suggested beamformer aims to maximize the array directivity factor while maintaining fixed beamwidth over a wide range of frequencies. The introduced methodology simultaneously selects the ring placements and designs the beamformer weights that achieve optimal performance. We present the problem constraints and cost function, resulting in a sparse beamformer design. In addition, a uniformly spaced beamformer is incorporated into the optimization cost function to attain improved performance. Time-domain implementation of the ideal beamformer is also presented in this work. Furthermore, in the proposed array configuration, all sensors on each ring share the same weight value. Thus, the computational complexity and the physical hardware required in a physical setup are significantly reduced. Experimental results demonstrate the advantages of the nonuniform beamformer compared to a uniform beamformer in terms of directivity index, white noise gain, and sidelobe attenuation.
Avital Kleiman, Israel Cohen, Baruch Berdugo
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Convolutional Sparse Coding Fast Approximation With Application to Seismic Reflectivity Estimation
abstract
In sparse coding, we attempt to extract features of input vectors, assuming that the data is inherently structured as a sparse superposition of basic building blocks. Similarly, neural networks perform a given task by learning features of the training dataset. Recently, both data- and model-driven feature extracting methods have become extremely popular and have achieved remarkable results. Nevertheless, practical implementations are often too slow to be employed in real-life scenarios, especially for real-time applications. We propose a speed-up upgraded version of the classic iterative thresholding algorithm (ITA), which produces a good approximation of the convolutional sparse code (CSC) within 2–5 iterations. The speed advantage is gained mostly from the observation that most solvers are slowed down by inefficient global thresholding. The main idea is to normalize each data point by the local receptive field energy, before applying a threshold. This way, the natural inclination toward strong feature expressions is suppressed, so that one can rely on a global threshold that can be easily approximated, or learned during training. The proposed algorithm can be employed with a known predetermined dictionary, or with a trained dictionary. The trained version is implemented as a neural net designed as the unfolding of the proposed solver. The performance of the proposed solution is demonstrated via the seismic inversion problem in both synthetic and real data scenarios. We also provide theoretical guarantees for a stable support recovery, namely we prove that under certain conditions, the true support is perfectly recovered within the first iteration.
Deborah Pereg, Israel Cohen, Anthony A. Vassiliou
IEEE Trans. Geosci. Remote. Sens.2
2021 Combined Differential Beamforming With Uniform Linear Microphone Arrays
abstract
While differential beamformers have been widely used in voice communication and human-machine speech interface systems to enhance speech signals of interest, how to design such beamformers that on the one hand can achieve the highest possible directivity factor (DF) and on the other hand are able to obtain a certain level of white noise gain (WNG), so that they are robust enough to sensors’ self noise and array imperfections is still a challenging issue. This paper studies the problem of robust differential beamforming with small-size arrays to achieve a high DF. It presents a method for the design of differential beamformers with uniform linear arrays. We first generate differential pressure signals by applying the recently developed forward spatial difference operator to the outputs of the array with pressure sensors. The pressure microphone observation signals and the differential pressure signals are then put together, and a combined beamformer is subsequently designed, which consists of two subbeamformers, one operates on the pressure microphone observations and the other on the differential pressure signals. A new class of combined differential beamformers are introduced, which can achieve different levels of compromises between DF and WNG using an adjustable parameter.
Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen
ICASSP4
2021 Deep Residual Echo Suppression With A Tunable Tradeoff Between Signal Distortion And Echo Suppression
abstract
In this paper, we propose a residual echo suppression method using a UNet neural network that directly maps the outputs of a linear acoustic echo canceler to the desired signal in the spectral domain. This system embeds a design parameter that allows a tunable tradeoff between the desired-signal distortion and residual echo suppression in double-talk scenarios. The system employs 136 thousand parameters, and requires 1.6 Giga floating-point operations per second and 10 Mega-bytes of memory. The implementation satisfies both the timing requirements of the AEC challenge and the computational and memory limitations of on-device applications. Experiments are conducted with 161 h of data from the AEC challenge database and from real independent recordings. We demonstrate the performance of the proposed system in real-life conditions and compare it with two competing methods regarding echo suppression and desired-signal distortion, generalization to various environments, and robustness to high echo levels.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP2
2021 Window Beamformer for Sparse Concentric Circular Array
abstract
This work proposes a practical fixed beamformer for a sparse concentric circular array (CCA) of microphones. Firstly, a method is proposed which enables the reduction of the number of microphones in a CCA based on the optimal placement of the microphones to maximize the directivity-factor (DF). Thereafter, a beamformer is derived which strives to obtain a beamwidth that is not smaller than a desired target beamwidth (in both the elevation and azimuth) without neglecting the DF and white-noise-gain (WNG) across the speech spectrum. The proposed beamformer weights the microphones based on their corresponding ring number or radius, and further uses a Gaussian-window function to emphasize the microphones which are aligned with the direction-of-arrival of the signal. The results showcase the superiority of the proposed Gaussian-window (GW) beamformer over known solutions.
Rajib Sharma, Israel Cohen, Baruch Berdugo
ICASSP2
2021 Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone Arrays
abstract
Differential beamformers with concentric circular microphone arrays (CCMAs) are desirable for use in various applications since they can form frequency-invariant spatial responses, have better beam steering flexibility than linear arrays, and suffer less with beampattern irregularity and white noise amplification than circular microphone arrays (CMAs). The methods developed previously for differential beamforming with CCMAs are based on the series expansion. Such methods need to know the analytic form of the target beam-pattern, which may not be accessible in practice. Furthermore, expansion error may lead to erroneous solution, which can cause noise amplification instead of reduction. In this paper, we extend our recently developed beamforming method for CMAs to the design of differential beamformers with CCMAs, which takes advantage of the symmetric null constraints from the beampattern. Simulations are performed to justify the properties of the proposed approach.
Xuehan Wang, Gongping Huang, Israel Cohen, Jacob Benesty, Jingdong Chen
ICASSP3
2021 On the Design of Square Differential Microphone Arrays with a Multistage Structure
abstract
This paper studies the problem of designing square differential microphone arrays (SDMAs). It presents a multistage approach, which first divides an SDMA composed of M2microphones into (M − 1)2subarrays with each subarray being a 2 × 2 square array formed by four adjacent microphones. Then, differential beamforming is performed with each subarray in the first-stage. The first-stage differential beamformers’ outputs are subsequently used as the inputs of the second stage to form (M − 2)2subarrays and a second-stage differential beamforming is then performed. Continuing this process till the (M −1)th stage, we obtain the final output of the SDMA. The SDMA designed in such a multistage structure has two important properties. First, the global weighting matrix is equal to the two dimensional convolution of weighting matrices from the first stage to the last one. Second, the global beampattern is equal to the product of beampatterns from all stages. Consequently, we can combine different kinds of beamformers in different stages and have better control of the performance metrics.
Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen
ICASSP5
2021 Nonlinear Acoustic Echo Cancellation with Deep Learning
abstract
We propose a nonlinear acoustic echo cancellation system, which aims to model the echo path from the far-end signal to the near-end microphone in two parts. Inspired by the physical behavior of modern hands-free devices, we first introduce a novel neural network architecture that is specifically designed to model the nonlinear distortions these devices induce between receiving and playing the far-end signal. To account for variations between devices, we construct this network with trainable memory length and nonlinear activation functions that are not parameterized in advance, but are rather optimized during the training stage using the training data. Second, the network is succeeded by a standard adaptive linear filter that constantly tracks the echo path between the loudspeaker output and the microphone. During training, the network and filter are jointly optimized to learn the network parameters. This system requires 17 thousand parameters that consume 500 Million floating-point operations per second and 40 Kilo-bytes of memory. It also satisfies hands-free communication timing requirements on a standard neural processor, which renders it adequate for embedding on hands-free communication devices. Using 280 hours of real and synthetic data, experiments show advantageous performance compared to competing methods.
Amir Ivry, Israel Cohen, Baruch Berdugo
Interspeech2
2021 Adaptive line enhancer for nonstationary harmonic noise reduction
Aviva Atkins, Israel Cohen, Jacob Benesty
Comput. Speech Lang.2
2021 Time Difference of Arrival Estimation Based on a Kronecker Product Decomposition
abstract
Time difference of arrival (TDOA) estimation, which often serves as the fundamental step for a source localization or a beamforming system, has a significant practical importance in a wide spectrum of applications. To deal with reverberation, the TDOA estimation problem is often transformed into one of identifying the relative acoustic impulse responses. This letter presents a method to efficiently identify the relative acoustic impulse response between two microphones for TDOA estimation based on the so-called Kronecker product decomposition. By decomposing the relative impulse response into a series of Kronecker products of shorter filters, the original channel identification problem with a long impulse response is converted into one of identifying a number of short filters. Since the TDOA information is embedded only in the direct path of the relative impulse response, the dimension of the Kronecker product decomposition can be very small and, as a result, the developed algorithm is expected to work well in real environments with a small number of data snapshots.
Xianrui Wang, Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen
IEEE Signal Process. Lett.5
2021 Robust Dereverberation With Kronecker Product Based Multichannel Linear Prediction
abstract
Reverberation impairs not only the speech quality, but also intelligibility. The weighted-prediction-error (WPE) method, which estimates the late reverberation component based on a multichannel linear predictor, is by far one of the most effective algorithms for dereverberation. Generally, the WPE prediction filter in every short-time-Fourier-transform (STFT) subband has to be long enough to estimate accurately the late reverberation component. As a consequence, WPE is computationally expensive, which makes it difficult to implement into real-time embedded or edge computing devices. Moreover, WPE is sensitive to additive noise and its performance may suffer from dramatic degradation even in environments where the signal-to-noise ratio (SNR) is high. To address these drawbacks, this letter proposes to decompose the multichannel linear prediction filter as a Kronecker product of a temporal (interframe) prediction filter and a spatial filter. An iterative algorithm is then developed to optimize the two filters. In comparison with the original WPE algorithm, the presented method not only exhibits better performance in terms of dereverberation and robustness to additive noise, as there are fewer parameters to estimate for a given number of observation signal samples, but is also computationally more efficient, since the dimensions of the covariance matrices after Kronecker product decomposition are smaller.
Wenxing Yang, Gongping Huang, Jingdong Chen, Jacob Benesty, Israel Cohen, Walter Kellermann
IEEE Signal Process. Lett.5
2021 On the Design of Differential Kronecker Product Beamformers
abstract
In this paper, we present a generalized approach for differential microphone array (DMA) beamforming in the short-time Fourier transform (STFT) domain. We propose a multistage beamforming approach, which considers a Kronecker product (KP) decomposition of the global beamformer into two independent sub-beamformers. We derive differential KP beamformers according to different criteria and analyze their performances, which are tuned by three design parameters. These parameters allow a high beamforming design flexibility; in particular, non-differential or non-KP beamformers may be obtained as special cases. Depending on the selection of parameters, we demonstrate a preferable performance with the new approach with respect to the white noise gain and directivity factor measures. In addition, we consider the task of speech enhancement. We show that differential KP beamformers perform better than non-differential and non-KP beamformers in terms of the quality and intelligibility of their respective time-domain enhanced signals, particularly in moderately reverberant environments.
Gal Itzhak, Jacob Benesty, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Steering Study of Linear Differential Microphone Arrays
abstract
Differential microphone arrays (DMAs) can achieve high directivity and frequency-invariant spatial response with small apertures; they also have a great potential to be used in a wide spectrum of applications for high-fidelity sound acquisition. Although many efforts have been made to address the design of linear DMAs (LDMAs), most developed methods so far only work for the situation where the source of interest is incident from the endfire direction. This paper studies the steering problem of differential beamformers with linear microphone arrays. We present new insights into beam steering of LDMAs and propose a series of steerable differential beamformers. The major contributions of this paper are as follows. 1) A series of ideal functions are defined to describe the ideal, target beampatterns of LDMAs. 2) We prove that first-order differential beamformers with linear microphone arrays are not steerable and their mainlobes can only be at the endfire directions. 3) We deduce the fundamental conditions for designing steerable differential beamformers with LDMAs. 4) We develop a method to design steerable beamformers with LDMAs using null constraints. Simulations and experiments validate the properties of the developed method.
Jilu Jin, Gongping Huang, Xuehan Wang, Jingdong Chen, Jacob Benesty, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Controlling Elevation and Azimuth Beamwidths With Concentric Circular Microphone Arrays
abstract
Solutions for frequency-invariant or constant-beamwidth beamforming with concentric circular arrays (CCAs) generally control only the azimuth-beampattern. This work introduces a beamforming methodology, which instead of designing a fixed beamwidth across the spectrum, maximizes the directivity-factor under the constraint that the beamwidth is greater than or equal to a pre-defined target-beamwidth, in both the elevation and the azimuth. The principal idea lies in dynamically weighting the microphones of the CCA based on two factors - (i) their corresponding ring number or radius, and (ii) their alignment with the signal-path. The parameters controlling the weights are determined using a modified gradient-descent algorithm. Experimental results show the flexibility and advantage of the proposed methodology over state-of-the-art methods.
Rajib Sharma, Israel Cohen, Baruch Berdugo
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Beamforming with Cube Microphone Arrays Via Kronecker Product Decompositions
abstract
Microphone arrays combined with beamforming have been widely used to solve many important acoustic problems in a wide range of applications. Much effort has been devoted in the literature to microphone array beamforming, among which the Kronecker product beamforming method developed recently has demonstrated some interesting properties. Generally, this method decomposes the global beamforming filter into a Kronecker product of a number of sub-beamforming filters, each of which corresponds to a virtual subarray and can be designed individually. This decomposition not only reduces significantly the number of beamforming coefficients, but also can be explored to improve the robustness and flexibility of beamforming. This paper extends Kronecker product beamforming from two-dimensional arrays into three-dimensional cube arrays. We consider two decompositions, i.e., fully and partially separable ones. The former decomposes the entire array into three linear subarrays while the latter decomposes the entire array into a linear subarray and a planar one. Then, for each case, we derive the Kronecker product maximum white noise gain beamformer, the Kronecker product approximate maximum directivity factor (DF) beamformer, the Kronecker product null-steering beamformer, and the Kronecker product iterative maximum DF beamformer. Simulation results demonstrate the properties and advantages of the proposed beamformers.
Xuehan Wang, Jacob Benesty, Jingdong Chen, Gongping Huang, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Greedy Sparse Array Design for Optimal Localization under Spatially Prioritized Source Distribution
abstract
A common approach for acoustic source localization is based on finding the maximum of a spatial cost function, such as the steered response power (SRP) function. The shape of the SRP highly depends on the constellation of sensors within the array layout, and have a direct impact on the performance. Thus, an array may be designed to produce high localization performance and small error regions, especially when a spatially prioritized source location distribution is taken into account. We introduce a new measure called power spread, which quantifies the localization error region. Then, we propose a greedy algorithm for a sparse array design, aiming to minimize the power spread for optimal error region in a given area of interest. Simulations demonstrate that the proposed design, compared to standard linear array design and random array design, obtains superior performance in terms of power spread and localization error, with a reasonable computational effort.
Yotam Gershon, Yaakov Buchris, Israel Cohen
ICASSP3
2020 Robust and steerable kronecker product differential beamforming With rectangular microphone arrays
abstract
Differential microphone arrays (DMAs), a class of welldesigned small-size arrays combined with differential beamforming, are very useful for processing broadband acoustic, audio, and speech signals in a wide range of applications. However, most efforts in the literature so far have been devoted to linear, circular, and spherical arrays. In this paper, we consider rectangular shapes of planar microphone arrays. Instead of adopting the traditional differential beamforming methods developed in the literature, we present a differential beamforming method based on the so-called Kronecker product. We first decompose the entire rectangular array into two virtual rectangular sub-arrays so that the steering vector of the entire array is the Kronecker product of the steering vectors of the two smaller virtual rectangular sub-arrays. We use the first virtual rectangular array, which is much smaller in size than the entire array but well satisfies the basic requirements for differential beamforming, to design a steerable differential beamformer. For the second virtual rectangular array, we can design either the delay-and-sum (DS) beamformer, which helps to improve the robustness of the global differential beamformer, or an adaptive beamformer, which makes the global differential beamformer adaptive. This method has many interesting properties, particularly the designed beamformer is fully steerable, and its robustness and the array gain can be easily controlled.
Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen
ICASSP4
2020 Evaluation of Deep-Learning-Based Voice Activity Detectors and Room Impulse Response Models in Reverberant Environments
abstract
State-of-the-art deep-learning-based voice activity detectors (VADs) are often trained with anechoic data. However, real acoustic environments are generally reverberant, which causes the performance to significantly deteriorate. To mitigate this mismatch between training data and real data, we simulate an augmented training set that contains nearly five million utterances. This extension comprises of anechoic utterances and their reverberant modifications, generated by convolutions of the anechoic utterances with a variety of room impulse responses (RIRs). We consider five different models to generate RIRs, and five different VADs that are trained with the augmented training set. We test all trained systems in three different real reverberant environments. Experimental results show 20% increase on average in accuracy, precision and recall for all detectors and response models, compared to anechoic training. Furthermore, one of the RIR models consistently yields better performance than the other models, for all the tested VADs. Additionally, one of the VADs consistently outperformed the other VADs in all experiments.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP2
2020 Adaptive and hybrid Kronecker product beamforming for far-field speech signals
Rajib Sharma, Israel Cohen, Jacob Benesty
Speech Commun.2
2020 Beamforming With Small-Spacing Microphone Arrays Using Constrained/Generalized LASSO
abstract
In this letter, we develop an approach to the design of beamformers with small-spacing uniform linear microphone arrays by incorporating sparseness constraints for attenuating scattered interference incident from some pre-specified ranges of directions of arrival. The design process is formulated as a constrained LASSO problem. By adjusting the value of a tuning parameter, the proposed method can make compromises among three important yet conflicting (especially at low frequencies) performance measures of small-spacing microphone arrays, i.e., the directivity factor (DF), which quantifies the array spatial gain, the white noise gain (WNG), which evaluates the robustness of the beamformer, and the signal-to-interference-ratio (SIR) gain with respect to scattered interference. Simulation results illustrate the properties of the developed approach.
Xianghui Wang, Jacob Benesty, Jingdong Chen, Israel Cohen
IEEE Signal Process. Lett.4
2020 Joint Sparse Concentric Array Design for Frequency and Rotationally Invariant Beampattern
abstract
Frequency-invariant concentric arrays are fundamental components in some real-world applications, like teleconferencing, voice service devices, underwater acoustics, and others, where the azimuthal arrival direction of the desired signal is varying. The fact that the demand for limited hardware and computational resources in such applications is essential, motivates the use of a sparse design which can optimize both the number of the required sensors and the complex weights of the beamformer. Herein, we propose a new greedy based joint-sparse design of frequency and rotationally invariant concentric arrays which preserves the properties of the designed directivity pattern for different azimuthal directions of steering. Simulation results show that the greedy sparse design, compared to uniform and random designs, gives superior performance in terms of array gain, and frequency and rotationally invariant beampattern, with a reasonable computational and hardware resources.
Yaakov Buchris, Israel Cohen, Jacob Benesty, Alon Amar
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Differential Beamforming on Graphs
abstract
We study differential beamforming from a graph perspective. The microphone array used for differential beamforming is viewed as a graph, where its sensors correspond to the nodes, the number of microphones corresponds to the order of the graph, and linear spatial difference equations among microphones are related to graph edges. Specifically, for the first-order differential beamforming with an array of M microphones, each pair of adjacent microphones are directly connected, resulting in M - 1 spatial difference equations. On a graph, each of these equations corresponds to a 2-clique. For the second-order differential beamforming, each three adjacent microphones are directly connected, resulting in M - 2 second-order spatial difference equations, and each of these equations corresponds to a 3-clique. In an analogous manner, the differential microphone array for any order-of-differential beamforming can be viewed as a graph. From this perspective, we then derive a class of differential beamformers, including the maximum white noise gain beamformer, the maximum directivity factor one, and optimal compromising beamformers. Simulations are presented to demonstrate the performance of the derived differential beamformers.
Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 A Simple Theory and New Method of Differential Beamforming With Uniform Linear Microphone Arrays
abstract
This article presents a theoretical study of differential beamforming with uniform linear arrays. By defining a forward spatial difference operator, any order of the spatial difference of the observed signals can be represented as a product of a difference operator matrix and the microphone array observations. Consequently, differential beamforming is implemented in two stages, where the first one obtains spatial difference of the observations and the second stage optimizes the beamformer. The major contributions of this article are as follows. First, we propose a new theory of differential beamforming with uniform linear arrays, which shows clearly the connection between the conventional differential beamforming and the null-constrained differential beamforming methods. This provides some new insight into the design of differential beamformers. Second, we deduce some new differential beamformers, where conventional beamforming may be seen as a particular case. Specifically, we derive the maximum white noise gain (MWNG), maximum directivity factor (MDF), parameterized MDF, and parameterized maximum front-to-back ratio differential beamformers. Third, we further extend the idea of how to design optimal differential beamformers by combining both the observed signals and their spatial differences.
Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 On the Design of Flexible Kronecker Product Beamformers with Linear Microphone Arrays
abstract
This paper proposes a method for the design of flexible Kronecker product beamformers based on the decomposition of the steering vector of a physical array as a Kronecker product of steering vectors of two smaller virtual arrays. With this decomposition, the global beamforming filter is designed by optimizing the two sub-beamformers in a cascaded manner, which can offer much flexibility to control the performance of beamforming or control the compromise between different, conflicted performance measures. In comparison with a recently developed method that restricts the number of microphones of the given physical array to a multiplication of two integers, each corresponding to the number of sensors of one virtual array, the approach in this work decomposes the physical array in such a way that the sensors in the two virtual arrays may share positions and the number of microphones of the physical array can be any positive integer. Simulations demonstrate the properties of the proposed approach.
Wenxing Yang, Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen
ICASSP4
2019 Three-dimensional sparse seismic deconvolution based on earth Q model
Deborah Pereg, Israel Cohen, Anthony A. Vassiliou
Signal Process.2
2019 Nonlinear Kronecker product filtering for multichannel noise reduction
Gal Itzhak, Jacob Benesty, Israel Cohen
Speech Commun.3
2019 Incoherent Synthesis of Sparse Arrays for Frequency-Invariant Beamforming
abstract
Frequency-invariant beamformers are used to prevent signal waveform distortions in real world applications like audio, underwater acoustics, and radar. Most of existing methods assume uniform arrays, and only few consider sparse designs, which may lead to higher performance in terms of robustness and directivity factor. We propose an incoherent approach that first determines for each frequency bin a sparse set of sensors positions. Subsequently, by using tools of dimensionality reduction and clustering, these selections are merged together yielding the optimal sensors on a sparse array layout. We present design examples of sparse linear and planar superdirective array designs. We show that the proposed incoherent sparse design obtains superior performance in terms of white noise gain, directivity factor, and computational load compared to a uniform array design and compared to a coherent sparse approach, where the sensors' locations and the beamformer coefficients are optimized simultaneously for all frequencies.
Yaakov Buchris, Alon Amar, Jacob Benesty, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Differential Kronecker Product Beamforming
abstract
Differential beamformers have attracted much interest over the past few decades. In this paper, we introduce differential Kronecker product beamformers that exploit the structure of the steering vector to perform beamforming differently from the well-known and studied conventional approach. We consider a class of microphone arrays that enable to decompose the steering vector as a Kronecker product of two steering vectors of smaller virtual arrays. In the proposed approach, instead of directly designing the differential beamformer, we break it down following the decomposition of the steering vector, and show how to derive differential beamformers using the Kronecker product formulation. As demonstrated, the Kronecker product decomposition facilitates further flexibility in the design of differential beamformers and in the tradeoff control between the directivity factor and the white noise gain.
Israel Cohen, Jacob Benesty, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 On Robust and High Directive Beamforming With Small-Spacing Microphone Arrays for Scattered Sources
abstract
This paper is devoted to beamforming with small-spacing microphone arrays for processing broadband and scattered acoustic sources. It presents a maximum diffuse noise gain (MDNG) beamformer in this context using the joint diagonalization technique, which is effective in suppressing diffuse and directional noise, but at a price of low white noise gain (WNG). We also introduce a maximum WNG (MWNG) beamformer, which is robust to the array imperfections, but paying a price of sacrificing the diffuse noise gain (DNG). To make a tradeoff between WNG and DNG so that the beamformer, on the one hand, can achieve high directivity and, on the other hand, is robust to implement, we propose a generalized MDNG beamformer, which includes both the MDNG and MWNG beamformers as particular cases. Simulations are conducted to illustrate the properties and advantages of the proposed beamformers.
Xianghui Wang, Israel Cohen, Jingdong Chen, Jacob Benesty
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Multi-View Source Localization Based on Power Ratios
abstract
Despite attracting significant research efforts, the problem of source localization in noisy and reverberant environments remains challenging. Novel learning-based methods attempt to solve the problem by modelling the acoustic environment from the observed data. Typically, appropriate feature vectors are defined, and then used for constructing a model, which maps the extracted features to the corresponding source positions. In this paper, we focus on localizing a source using a distributed network with several arrays of unidirectional microphones. We introduce new feature vectors, which utilize the special characteristic of unidirectional microphones, receiving different parts of the reverberated speech. The new features are computed locally for each array, using the power-ratios between its measured signals, and are used to construct a local model, representing the unique view point of each array. The models of the different arrays, conveying distinct and complementing structures, are merged by a Multi-View Gaussian Process (MVGP), mapping the new features to their corresponding source positions. Based on this unifying model, a Bayesian estimator is derived, exploiting the relations conveyed by the covariance terms of the MVGP. The resulting localizer is shown to be robust to noise and reverberation, utilizing a computationally efficient feature extraction.
Bracha Laufer-Goldshtein, Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP3
2018 On Speech Enhancement Using Microphone Arrays in the Presence of Co-Directional Interference
abstract
Beamforming using microphone arrays has been widely used for enhancing speech signals of interest and suppressing noise and interference in a wide range of applications. In order to make it work, beamforming generally assumes that the speech source of interest and the interference source are incident to the array from different directions. In this paper, we study the case where both the speech and interference sources come from the same direction. A linearly constrained minimum variance (LCMV) beamformer is derived in this scenario based on the so-called widely linear (WL) estimation framework in the frequency domain. We analyze this beamformer and show how its performance depends on the second-order non-circularity of the desired speech and interference sources.
Xin Leng, Jingdong Chen, Jacob Benesty, Israel Cohen
ICASSP4
2018 A deep architecture for audio-visual voice activity detection in the presence of transients
Ido Ariav, David Dov, Israel Cohen
Signal Process.3
2018 Frequency-Domain Design of Asymmetric Circular Differential Microphone Arrays
abstract
Circular differential microphone arrays (CDMAs) facilitate compact superdirective beamformers whose beampatterns are nearly frequency invariant. In contrast to linear differential microphone arrays where the optimal steering direction is at the endfire, CDMAs provide perfect steering for all azimuthal directions. Herein, we extend the traditional symmetric model of DMAs and establish an analytical asymmetric model for Nth-order CDMAs. This model exploits the circular geometry to eliminate the inherent limitation of symmetric beampatterns associated with a linear geometry and allows also asymmetric beampatterns. This new model is then used to develop asymmetric versions of two optimal commonly used beampatterns namely the hypercardioid and the supercardioid. Experimental results demonstrate the advantages of the asymmetric model compared to the traditional symmetric one, when additional directional constraints are imposed. The proposed model yields superior performance in terms of white noise gain, directivity factor, and front-to-back ratio, as well as more flexible design of nulls for the interfering signals.
Yaakov Buchris, Israel Cohen, Jacob Benesty
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Eigenvalue decomposition based estimators of carrier frequency offset in multicarrier underwater acoustic communication
abstract
We propose computationally efficient carrier frequency offset estimators for multicarrier underwater acoustic communication using identical pilot tones equi-spaced in the frequency domain. The first estimator uses the phase of the maximal eigenvector of a channel-dependent correlation matrix. Next, the phase of the minimal eigenvector of a channel-independent correlation matrix is combined with the first estimation using a weighted linear least squares principle. The third estimator solves a generalized eigenvalue decomposition problem by jointly considering the two correlation matrices, and then performs a similar second step as the previous estimator. Simulations and pool trials show that the proposed estimators achieve similar performance as common estimation techniques while surpassing them in severe environments.
Gilad Avrashi, Alon Amar, Israel Cohen, Milica Stojanovic
ICASSP3
2017 Iterative diffusion-based anomaly detection
abstract
Diffusion maps, when applied to large datasets, are typically constructed by a process of sampling and out-of-sample function extension. However, the performance of anomaly detection in large data when using diffusion maps is sensitive to the chosen samples. In this paper we propose an iterative data-driven approach to improve the sample set and diffusion maps representation. By updating the sample set with suspicious points detected in the previous iteration, the constructed diffusion maps better separate the anomaly from the normal points in each iteration. Experimental results in side-scan sonar images demonstrate the improvement gained by our iterative sampling compared to random sampling and other competing detection algorithms.
Gal Mishne, Israel Cohen
ICASSP2
2017 Multichannel Semi-blind Deconvolution (MSBD) of seismic signals
Merabi Mirel, Israel Cohen
Signal Process.2
2017 Seismic signal recovery based on Earth Q model
Deborah Pereg, Israel Cohen
Signal Process.2
2017 FIR-based symmetrical acoustic beamformer with a constant beamwidth
Oren Rosen, Israel Cohen, David Malah
Signal Process.2
2017 Multimodal Kernel Method for Activity Detection of Sound Sources
abstract
We consider the problem of acoustic scene analysis of multiple sound sources. In our setting, the sound sources are measured by a single microphone, and a particular source of interest is also captured by a video camera during a short time interval. The goal in this paper is to detect the activity of the source of interest even when the video data are missing, while ignoring the other sound sources. To address this problem, we propose a kernel-based algorithm that incorporates the audio-visual data by a combination of affinity kernels, constructed separately from the audio and the video data. We introduce a distance measure between data points that is associated with the source of interest, while reducing the effect of the other (interfering) sources. Using this distance, we devise a measure for the presence of the source of interest, which is naturally extended to time intervals, in which only the audio signal is available. Experimental results demonstrate the improved performance of the proposed algorithm compared to competing approaches implying the significance of the video signal in the analysis of complex acoustic scenes.
David Dov, Ronen Talmon, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and Korean
abstract
In studies on artificial bandwidth extension (ABE), there is a lack of international coordination in subjective tests between multiple methods and languages. Here we present the design of absolute category rating listening tests evaluating 12 ABE variants of six approaches in multiple languages, namely in American English, Chinese, German, and Korean. Since the number of ABE variants caused a higher-than-recommended length of the listening test, ABE variants were distributed into two separate listening tests per language. The paper focuses on the listening test design, which aimed at merging the subjective scores of both tests and thus allows for a joint analysis of all ABE variants under test at once. A language-dependent analysis, evaluating ABE variants in the context of the underlying coded narrowband speech condition showed statistical significant improvement in English, German, and Korean for some ABE solutions.
Johannes Abel, Magdalena Kaniewska, Cyril Guillaume, Wouter Tirry, Hannu Pulakka, Ville Myllylä, Jari Sjoberg, Paavo Alku, Itai Katsir, David Malah, Israel Cohen, M. A. Tugtekin Turan, Engin Erzin, Thomas Schlien, Peter Vary, Amr H. Nour-Eldin, Peter Kabal, Tim Fingscheidt
ICASSP11
2016 Kernel Method for Voice Activity Detection in the Presence of Transients
abstract
Voice activity detection in the presence of transient interferences is a challenging problem since transients are often detected incorrectly as speech by existing detectors. In this paper, we deviate from traditional approaches and take a geometric standpoint, in which the key element in obtaining an accurate voice activity detection is finding a metric that appropriately distinguishes between speech and transients. For example, speech and transients may often appear similar through the Euclidean distance when represented, e.g., by the Mel-frequency cepstral coefficients, thereby resulting in incorrect speech detection. To address this challenge, we propose to use a metric based on the statistics of the signal in short temporal windows and justify its use by modeling speech and transients by their latent generating variables. These latent variables may be related to physical constraints controlling the generation of the signal, and, as such, they accurately represent the content of the signal - speech or transient. We show that the Euclidean distance between the latent variables is approximated by the proposed metric. Then, by incorporating this metric into a kernel-based manifold learning method, we devise a measure of voice activity and show it leads to improved detection scores compared with competing detectors.
David Dov, Ronen Talmon, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Musical key extraction using diffusion maps
Ofir Lindenbaum, Arie Yeredor, Israel Cohen
Signal Process.3
2015 Out-of-sample extension of band-limited functions on homogeneous manifolds using diffusion maps
Saman Mousazadeh, Israel Cohen
Signal Process.2
2015 Embedding and function extension on directed graph
Saman Mousazadeh, Israel Cohen
Signal Process.2
2015 Combined Beamformers for Robust Broadband Regularized Superdirective Beamforming
abstract
Superdirective fixed beamformers are known to attain high directivity factors, but are extremely sensitive to uncorrelated noise and slight errors in the array elements, which are modeled by the beamformer white noise gain measure. The delay-and-sum beamformer, on the other hand, manages to maximize the white noise gain, but suffers from a very low directivity factor. In this paper, we discuss the design of a broadband beamformer which controls both the directivity factor and the white noise gain. We combine a regularized version of the superdirective beamformer together with the delay-and-sum beamformer to create a robust regularized superdirective beamformer. We derive analytic closed-form expressions of the beamformer gain responses, and extend them to derive a beamformer with full control of the desired white noise gain or the directivity factor. The proposed approach offers a simple and robust broadband beamformer with controllable characteristics, shown here through persuasive simulation results.
Reuven Berkun, Israel Cohen, Jacob Benesty
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Audio-Visual Voice Activity Detection Using Diffusion Maps
abstract
The performance of traditional voice activity detectors significantly deteriorates in the presence of highly nonstationary noise and transient interferences. One solution is to incorporate a video signal which is invariant to the acoustic environment. Although several voice activity detectors based on the video signal were recently presented, merely few detectors which are based on both the audio and the video signals exist in the literature to date. In this paper, we present an audio-visual voice activity detector and show that the incorporation of both audio and video signals is highly beneficial for voice activity detection. The algorithm is based on a supervised learning procedure, and a labeled training data set is considered. The algorithm comprises a feature extraction procedure, where the features are designed to separate speech from nonspeech frames. Diffusion maps is applied separately and similarly to the features of each modality and builds a low dimensional representation. Using the new representation, we propose a measure for voice activity which is based on a supervised learning procedure and the variability between adjacent frames in time. The measures of the two modalities are merged to provide voice activity detection based on both the audio and the video signals. Experimental results demonstrate the improved performance of the proposed algorithm compared to state-of-the-art detectors.
David Dov, Ronen Talmon, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Graph-Based Supervised Automatic Target Detection
abstract
In this paper, we propose a detection method based on data-driven target modeling, which implicitly handles variations in the target appearance. Given a training set of images of the target, our approach constructs models based on local neighborhoods within the training set. We present a new metric using these models and show that, by controlling the notion of locality within the training set, this metric is invariant to perturbations in the appearance of the target. Using this metric in a supervised graph framework, we construct a low-dimensional embedding of test images. Then, a detection score based on the embedding determines the presence of a target in each image. The method is applied to a data set of side-scan sonar images and achieves impressive results in the detection of sea mines. The proposed framework is general and can be applied to different target detection problems in a broad range of signals.
Gal Mishne, Ronen Talmon, Israel Cohen
IEEE Trans. Geosci. Remote. Sens.3
2014 Multiscale anomaly detection using diffusion maps and saliency score
abstract
Recently, we presented a multiscale approach to anomaly detection in images, combining diffusion maps for dimensionality reduction and a nearest-neighbor-based anomaly score in the reduced dimension. When applying diffusion maps to images, usually a process of sampling and out-of-sample extension is used, which has limitations in regards to anomaly detection. To overcome the limitations, a multiscale approach was proposed, which drives the sampling process to ensure separability of the anomaly from the background clutter. In this paper, we propose a new anomaly score used in the diffusion map space, which shows increased performance. We show that this algorithm enables improved detection when tested on side-scan sonar images of sea-mines and compare it with competing algorithms.
Gal Mishne, Israel Cohen
ICASSP2
2014 Two dimensional noncausal AR-ARCH model: Stationary conditions, parameter estimation and its application to anomaly detection
Saman Mousazadeh, Israel Cohen
Signal Process.2
2014 Facial Image Compression using Patch-Ordering-Based Adaptive Wavelet Transform
abstract
Compression of frontal facial images is an appealing and important application. Recent work has shown that specially tailored algorithms for this task can lead to performance far exceeding JPEG2000. This letter proposes a novel such compression algorithm, exploiting our recently developed redundant tree-based wavelet transform. Originally meant for functions defined on graphs and cloud of points, this new transform has been shown to be highly effective as an image adaptive redundant and multi-scale decomposition. The key concept behind this method is reordering of the image pixels so as to form a highly smooth 1D signal that can be sparsified by a regular wavelet. In this work we bring this image adaptive transform to the realm of compression of aligned frontal facial images. Given a training set of such images, the transform is designed to best sparsify the whole set using a common feature-ordering. Our compression scheme consists of sparse coding using the transform, followed by entropy coding of the obtained coefficients. The inverse transform and a post-processing stage are used to decode the compressed image. We demonstrate the performance of the proposed scheme and compare it to other competing algorithms.
Idan Ram, Israel Cohen, Michael Elad
IEEE Signal Process. Lett.2
2014 Patch-Ordering-Based Wavelet Frame and Its Use in Inverse Problems
abstract
In our previous work [1] we have introduced a redundant tree-based wavelet transform (RTBWT), originally designed to represent functions defined on high dimensional data clouds and graphs. We have further shown that RTBWT can be used as a highly effective image-adaptive redundant transform that operates on an image using orderings of its overlapped patches. The resulting transform is robust to corruptions in the image, and thus able to efficiently represent the unknown target image even when it is calculated from its corrupted version. In this paper, we utilize this redundant transform as a powerful sparsity-promoting regularizer in inverse problems in image processing. We show that the image representation obtained with this transform is a frame expansion, and derive the analysis and synthesis operators associated with it. We explore the use of this frame operators to image denoising and deblurring, and demonstrate in both these cases state-of-the-art results.
Idan Ram, Israel Cohen, Michael Elad
IEEE Trans. Image Process.2
2013 ARCH and GARCH parameter estimation in presence of additive noise using particle methods
abstract
In this paper, we propose a new method based on particle filters for maximum likelihood (ML) estimation of the parameters of autoregressive conditional heteroscedasticity (ARCH) and generalized autoregressive conditional heteroscedasticity (GARCH) models. Our method is based on gradient descend method and active set method for maximizing the likelihood function over parameters under stationarity constraints. The gradient of the likelihood function of observation given the parameters of the model, which is needed for gradient based optimization algorithm, is estimated using particle methods. Simulation results show the advantage of the proposed method over competing techniques.
Saman Mousazadeh, Israel Cohen
ICASSP2
2013 Image denoising using NL-means via smooth patch ordering
abstract
In our recent work we proposed an image denoising scheme based on reordering of the noisy image pixels to a one dimensional (1D) signal, and applying linear smoothing filters on it. This algorithm had two main limitations: It did not take advantage of the distances between the noisy image patches, which were used in the reordering process; and the smoothing filters required a separate training set to be learned from. In this work, we propose an image denoising algorithm, which applies similar permutations to the noisy image, but overcomes the above two shortcomings. We eliminate the need for learning filters by employing the nonlocal means (NL-means) algorithm. We estimate each pixel as a weighted average of noisy pixels in union of neighborhoods obtained from different global pixel permutations, where the weights are determined by distances between the patches. We show that the proposed scheme achieves results which are close to the state-of-the-art.
Idan Ram, Michael Elad, Israel Cohen
ICASSP3
2013 Dominant speaker identification for multipoint videoconferencing
Ilana Volfin, Israel Cohen
Comput. Speech Lang.2
2013 Distributed Multiple Constraints Generalized Sidelobe Canceler for Fully Connected Wireless Acoustic Sensor Networks
abstract
This paper proposes a distributed multiple constraints generalized sidelobe canceler (GSC) for speech enhancement in anN-node fully connected wireless acoustic sensor network (WASN) comprisingMmicrophones. Our algorithm is designed to operate in reverberant environments with constrained speakers (including both desired and competing speakers). Rather than broadcastingMmicrophone signals, a significant communication bandwidth reduction is obtained by performing local beamforming at the nodes, and utilizing only transmission channels. Each node processes its own microphone signals together with the N + P transmitted signals. The GSC-form implementation, by separating the constraints and the minimization, enables the adaptation of the BF during speech-absent time segments, and relaxes the requirement of other distributed LCMV based algorithms to re-estimate the sources RTFs after each iteration. We provide a full convergence proof of the proposed structure to the centralized GSC-beamformer (BF). An extensive experimental study of both narrowband and (wideband) speech signals verifíes the theoretical analysis.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.3
2013 Performance of the SDW-MWF With Randomly Located Microphones in a Reverberant Enclosure
abstract
Beamforming with wireless acoustic sensor networks (WASNs) has recently drawn the attention of the research community. As the number of microphones grows it is difficult, and in some applications impossible, to determine their layout beforehand. A common practice in analyzing the expected performance is to utilize statistical considerations. In the current contribution, we consider applying the speech distortion weighted multi-channel Wiener filter (SDW-MWF) to enhance a desired source propagating in a reverberant enclosure where the microphones are randomly located with a uniform distribution. Two noise fields are considered, namely, multiple coherent interference signals and a diffuse sound field. Utilizing the statistics of the acoustic transfer function (ATF), we derive a statistical model for two important criteria of the beamformer (BF): the signal to interference ratio (SIR), and the white noise gain. Moreover, we propose reliability functions, which determine the probability of the SIR and white noise gain to exceed a predefined level. We verify the proposed model with an extensive simulative study.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.3
2013 Voice Activity Detection in Presence of Transient Noise Using Spectral Clustering
abstract
Voice activity detection has attracted significant research efforts in the last two decades. Despite much progress in designing voice activity detectors, voice activity detection (VAD) in presence of transient noise is a challenging problem. In this paper, we develop a novel VAD algorithm based on spectral clustering methods. We propose a VAD technique which is a supervised learning algorithm. This algorithm divides the input signal into two separate clusters (i.e., speech presence and speech absence frames). We use labeled data in order to adjust the parameters of the kernel used in spectral clustering methods for computing the similarity matrix. The parameters obtained in the training stage together with the eigenvectors of the normalized Laplacian of the similarity matrix and Gaussian mixture model (GMM) are utilized to compute the likelihood ratio needed for voice activity detection. Simulation results demonstrate the advantage of the proposed method compared to conventional statistical model-based VAD algorithms in presence of transient noise.
Saman Mousazadeh, Israel Cohen
IEEE Trans. Speech Audio Process.2
2013 Single-Channel Transient Interference Suppression With Diffusion Maps
abstract
A transient is an abrupt or impulsive sound followed by decaying oscillations, e.g., keyboard typing and door knocking. Such sounds often arise as interference in everyday applications, e.g., hearing aids, hands-free accessories, mobile phones, and conference-room devices. In this paper, we present an algorithm for single-channel transient interference suppression. The main component of the proposed algorithm is the estimation of the spectral variance of the interference. We propose a statistical model of the transient interference and combine it with non-local filtering. We exploit the unique spectral structure of the transients along with their impulsive temporal nature to distinct them from speech. A particular attention is given to handling both short- and long-duration transients. Experimental results show that the proposed algorithm enables significant transient suppression for a variety of transient types.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.2
2013 Image Processing Using Smooth Ordering of its Patches
abstract
We propose an image processing scheme based on reordering of its patches. For a given corrupted image, we extract all patches with overlaps, refer to these as coordinates in high-dimensional space, and order them such that they are chained in the "shortest possible path," essentially solving the traveling salesman problem. The obtained ordering applied to the corrupted image implies a permutation of the image pixels to what should be a regular signal. This enables us to obtain good recovery of the clean image by applying relatively simple one-dimensional smoothing operations (such as filtering or interpolation) to the reordered set of pixels. We explore the use of the proposed approach to image denoising and inpainting, and show promising results in both cases.
Idan Ram, Michael Elad, Israel Cohen
IEEE Trans. Image Process.3
2012 A sparse blocking matrix for multiple constraints GSC beamformer
abstract
Modern high performance speech processing applications incorporate large microphone arrays. Complicated scenarios comprising multiple sources, motivate the use of the linearly constrained minimum variance (LCMV) beamformer (BF) and specifically its efficient generalized sidelobe canceler (GSC) implementation. The complexity of applying the GSC is dominated by the blocking matrix (BM). A common approach for constructing the BM is to use a projection matrix to the null-subspace of the constraints. The latter BM is denoted as the eigen-space BM, and requires M2complex multiplications, where M is the number of microphones. In the current contribution, a novel systematic scheme for constructing a multiple constraints sparse BM is presented. The sparsity of the proposed BM substantially reduces the complexity to K × (M - K) complex multiplications, where K is the number of constraints. A theoretical analysis of the signal leakage and of the blocking ability of the proposed sparse BM and of the eigen-space BM is derived. It is proven analytically, and tested for narrowband signals and for speech signals, that the blocking abilities of the sparse and of the eigen-space BMs are equivalent.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
ICASSP3
2012 Supervised system identification based on local PCA models
abstract
We propose a supervised system identification method for recovering an acoustic impulse response in a reverberant room. Unlike most existing methods, our algorithm is based on prior information given in the form of a training set of known impulse responses acquired in a controlled environment. By relying on the prior information, we train local Principal Component Analysis (PCA) models of impulse responses corresponding to several different regions in the room. We propose to crudely localize the respective source position, and subsequently, based on the appropriate local model, recover the impulse response. In order to approximate the source location, we introduce a specially-tailored distance measure which is based on an affinity between the trained local models. Experimental results in simulated noisy and reverberant environments demonstrate significant improvements over existing methods.
Tomer Koren, Ronen Talmon, Israel Cohen
ICASSP3
2012 Redundant Wavelets on Graphs and High Dimensional Data Clouds
abstract
In this paper, we propose a new redundant wavelet transform applicable to scalar functions defined on high dimensional coordinates, weighted graphs and networks. The proposed transform utilizes the distances between the given data points to construct tree-like structures. We modify the filter-bank decomposition scheme of the redundant wavelet transform by adding in each decomposition level operators that reorder the approximation coefficients. These reordering operators are derived by organizing the tree-node features so as to shorten the path that passes through these points. We explore the use of the proposed transform for the recovery of labels defined on point clouds and to image denoising, and show that in both cases the results are promising.
Idan Ram, Michael Elad, Israel Cohen
IEEE Signal Process. Lett.3
2012 Bayesian Focusing for Coherent Wideband Beamforming
abstract
In this paper, we present and study a Bayesian focusing transformation (BFT) for coherent wideband array processing, which takes into account the uncertainty of the direction of arrivals (DOAs). The Bayesian focusing method minimizes the mean-square error of the transformation over the probability density functions (pdfs) of the DOAs, thus achieving improved focusing accuracy over the entire bandwidth. In order to solve the Bayesian focusing problem, we derive and utilize a weighted extension of the wavefield interpolated narrowband generated subspace (WINGS) focusing transformation. We provide a closed-form expression for the optimal BFT and extend it to the case of directional sensors. We then consider a numerical computation scheme of the BFT in the angular domain. We show that if an angular sampling condition is satisfied then the angle domain approximation yields the optimal BFT. We also treat the important issue of robust focused minimum variance distortionless response (MVDR) beamformer. We analyze the sensitivity of the focused MVDR to the focusing errors and show that the array gain (AG) is inversely proportional to the square of the signal-to-noise ratio (SNR) for large values of the SNR, and highly sensitive to the focusing errors. In order to reduce this sensitivity we generalize the popular narrowband diagonal loaded MVDR to the focused wideband case, referred to as the Q-loaded focused MVDR wideband beamformer. We derive a closed-form analytic expression for the AG of the Q-loaded focused MVDR beamformer which depends on the focusing transformations. A numerical performance evaluation and simulations demonstrate the advantage of the BFT over that of other focusing transformations, for multiple source scenarios.
Yaakov Buchris, Israel Cohen, Miriam A. Doron
IEEE Trans. Speech Audio Process.2
2012 Supervised Graph-Based Processing for Sequential Transient Interference Suppression
abstract
In this paper, we present a supervised graph-based framework for sequential processing and employ it to the problem of transient interference suppression. Transients typically consist of an initial peak followed by decaying short-duration oscillations. Such sounds, e.g., keyboard typing and door knocking, often arise as an interference in everyday applications: hearing aids, hands-free accessories, mobile phones, and conference-room devices. We describe a graph construction using a noisy speech signal and training recordings of typical transients. The main idea is to capture the transient interference structure, which may emerge from the construction of the graph. The graph parametrization is then viewed as a data-driven model of the transients and utilized to define a filter that extracts the transients from noisy speech measurements. Unlike previous transient interference suppression studies, in this work the graph is constructed in advance from training recordings. Then, the graph is extended to newly acquired measurements, providing a sequential filtering framework of noisy speech.
Ronen Talmon, Israel Cohen, Sharon Gannot, Ronald R. Coifman
IEEE Trans. Speech Audio Process.2
2011 Performance analysis of a randomly spaced wireless microphone array
abstract
A randomly distributed microphone array is considered in this work. In many applications exact design of the array is impractical. The performance of these arrays, characterized by a large number of microphones deployed in vast areas, cannot be analyzed by traditional deterministic methods. We therefore derive a novel statistical model for performance analysis of the MWF beamformer. We consider the scenario of one desired source and one interfering source arriving from the far-field and impinging on a uniformly distributed linear array. A theoretical model for the MMSE is developed and verified by simulations. The applicability of the proposed statistical model for speech signals is discussed.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
ICASSP3
2011 Clustering and suppression of transient noise in speech signals using diffusion maps
abstract
Recently we have presented a novel approach for transient noise reduction that relies on non-local (NL) filtering. In this paper, we modify and extend our approach to support clustering and suppression of a few transient noise types simultaneously, by introducing two novel concepts. We observe that voiced speech spectral components are slowly varying compared to transient noise. Thus, by applying an algorithm for noise power spectral density (PSD) estimation, configured to track faster variations than pseudo-stationary noise, the PSD of speech components may be estimated. In addition, we utilize diffusion maps to embed the measurements into a new do main. We obtain a new representation which enables clustering of different transient noise types. The new representation is incorporated into a NL filter as a better affinity metric for averaging over transient instances. Experimental results show that the proposed algorithm enables clustering and suppression of multiple transient interferences.
Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP2
2011 AR-GARCH in Presence of Noise: Parameter Estimation and Its Application to Voice Activity Detection
abstract
This paper presents a new method for voice activity detection (VAD) based on the autoregressive-generalized autoregressive conditional heteroscedasticity (AR-GARCH) model. The speech signal is modeled as an AR-GARCH process in the time domain, and the likelihood ratio is computed and compared to a threshold. The time-varying variance of the speech signal needed for computing the likelihood function under speech presence hypothesis, is estimated using the AR-GARCH model. The model parameters are estimated using a novel technique based on the recursive maximum likelihood (RML) estimation. The variance of the additive noise, a critical issue in designing a VAD, is estimated using the improved minima controlled recursive averaging (IMCRA) method, which is properly modified to be applicable to noise variance estimation in the time domain. The performances of the VAD and the parameter estimation method are examined under several conditions. Experimental results indicate the robustness of the AR-GARCH based VAD both to noise variations and low signal-to-noise ratio (SNR) conditions.
Saman Mousazadeh, Israel Cohen
IEEE ACM Trans. Audio Speech Lang. Process.2
2011 Transient Noise Reduction Using Nonlocal Diffusion Filters
abstract
Enhancement of speech signals for hands-free communication systems has attracted significant research efforts in the last few decades. Still, many aspects and applications remain open and require further research. One of the important open problems is the single-channel transient noise reduction. In this paper, we present a novel approach for transient noise reduction that relies on non-local (NL) neighborhood filters. In particular, we propose an algorithm for the enhancement of a speech signal contaminated by repeating transient noise events. We assume that the time duration of each reoccurring transient event is relatively short compared to speech phonemes and model the speech source as an auto-regressive (AR) process. The proposed algorithm consists of two stages. In the first stage, we estimate the power spectral density (PSD) of the transient noise by employing a NL neighborhood filter. In the second stage, we utilize the optimally modified log spectral amplitude (OM-LSA) estimator for denoising the speech using the noise PSD estimate from the first stage. Based on a statistical model for the measurements and diffusion interpretation of NL filtering, we obtain further insight into the algorithm behavior. In particular, for given transient noise, we determine whether estimation of the noise PSD is feasible using our approach, how to properly set the algorithm parameters, and what is the expected performance of the algorithm. Experimental study shows good results in enhancing speech signals contaminated by transient noise, such as typical household noises, construction sounds, keyboard typing, and metronome clacks.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.2
2010 Frequency-based detector for improved resolvability of closely-spaced exponential signals
abstract
In this paper, we introduce a new local frequency-based detector for improved resolvability of closely-spaced exponential signals. The frequency resolution problem is defined as a hypothesis-testing problem in the frequency domain, and a corresponding detector, based on the generalized likelihood ratio test is constructed. We derive a theoretical resolution limit in terms of the signal-to-noise ratio, the observable-data length, and the probabilities of detection and false-alarm. Experimental results validate the theoretical results and demonstrate the effectiveness of the proposed detector over a time-based detector.
Yekutiel Avargel, Israel Cohen
ICASSP2
2010 Subspace tracking of multiple sources and its application to speakers extraction
abstract
In this paper we introduce a novel algorithm for extracting desired speech signals uttered by moving speakers contaminated by competing speakers and stationary noise in a reverberant environment. The proposed beamformer uses eigenvectors spanning the desired and interference signals subspaces. It relaxes the common requirement on the activity patterns of the various sources. A novel mechanism for tracking the desired and interferences subspaces is proposed, based on the projection approximation subspace tracking (deflation) (PASTd) procedure and on a union of subspaces procedure. This contribution extends previously proposed methods to deal with multiple speakers in dynamic scenarios.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
ICASSP3
2010 Particle filtering based recovery of noisy GARCH processes
abstract
In this paper, we address the problem of enhancement of a noisy GARCH process using a particle filter. We compare our approach experimentally to a previously developed recursive estimation scheme. Simulations indicate that a significant gain in performance is obtained, at the cost of higher sensitivity to errors in the GARCH parameters. The proposed method allows tackling arbitrary driving noise distributions as well as arbitrary fidelity criteria.
Tomer Michaeli, Israel Cohen
ICASSP2
2010 Speech enhancement in transient noise environment using diffusion filtering
abstract
Recently, we have presented a transient noise reduction algorithm for speech signals that relies on non-local diffusion filtering. By exploiting the repetitive nature of transient noises we proposed a simple and efficient algorithm, which enabled suppression of various noise types. In this paper, we incorporate a modified diffusion operator in order to obtain a more robust algorithm and further enhancement of the speech. We demonstrate the performance of the modified algorithm and compare it with a competing solution. We show that the proposed algorithm enables improved suppression of various transient interferences without any further computational burden.
Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP2
2010 Defect detection in patterned wafers using anisotropic kernels
Maria Zontak, Israel Cohen
Mach. Vis. Appl.2
2010 Monaural speech/music source separation using discrete energy separation algorithm
Yevgeni Litvin, Israel Cohen, Dan Chazan
Signal Process.2
2010 Simultaneous parameter estimation and state smoothing of complex GARCH process in the presence of additive noise
Saman Mousazadeh, Israel Cohen
Signal Process.2
2010 New Insights Into the MVDR Beamformer in Room Acoustics
abstract
The minimum variance distortionless response (MVDR) beamformer, also known as Capon's beamformer, is widely studied in the area of speech enhancement. The MVDR beamformer can be used for both speech dereverberation and noise reduction. This paper provides new insights into the MVDR beamformer. Specifically, the local and global behavior of the MVDR beamformer is analyzed and novel forms of the MVDR filter are derived and discussed. In earlier works it was observed that there is a tradeoff between the amount of speech dereverberation and noise reduction when the MVDR beamformer is used. Here, the tradeoff between speech dereverberation and noise reduction is analyzed thoroughly. The local and global behavior, as well as the tradeoff, is analyzed for different noise fields such as, for example, a mixture of coherent and non-coherent noise fields, entirely non-coherent noise fields and diffuse noise fields. It is shown that maximum noise reduction is achieved when the MVDR beamformer is used for noise reduction only. The amount of noise reduction that is sacrificed when complete dereverberation is required depends on the direct-to-reverberation ratio of the acoustic impulse response between the source and the reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction. When desiring both speech dereverberation and noise reduction, the results also demonstrate that the amount of noise reduction that is sacrificed decreases when the number of microphones increases.
Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot, Jacek Dmochowski
IEEE Trans. Speech Audio Process.3
2009 On a tradeoff between dereverberation and noise reduction using the MVDR beamformer
abstract
The minimum variance distortionless response (MVDR) beamformer can be used for both speech dereverberation and noise reduction. In this paper we analyse the tradeoff between the amount of speech dereverberation and noise reduction achieved by the MVDR beamformer. We show that the amount of noise reduction that is sacrificed when desiring both speech dereverberation and noise reduction depends on the direct-to-reverberation ratio of the acoustic transfer function between the desired source and a reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction.
Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot
ICASSP3
2009 Multichannel speech enhancement using convolutive transfer function approximation in reverberant environments
abstract
Recently, we have presented a transfer-function generalized sidelobe canceler (TF-GSC) beamformer in the short time Fourier transform domain, which relies on a convolutive transfer function approximation of relative transfer functions between distinct sensors. In this paper, we combine a delay-and-sum beamformer with the TF-GSC structure in order to suppress the speech signal reflections captured at the sensors in reverberant environments. We demonstrate the performance of the proposed beamformer and compare it with the TF-GSC. We show that the proposed algorithm enables suppression of reverberations and further noise reduction compared with the TF-GSC beamformer.
Ronen Talmon, Israel Cohen, Sharon Gannot
ICASSP2
2009 Defect detection in patterned wafers using multichannel Scanning Electron Microscope
Maria Zontak, Israel Cohen
Signal Process.2
2009 Late Reverberant Spectral Variance Estimation Based on a Statistical Model
abstract
In speech communication systems the received microphone signals are degraded by room reverberation and ambient noise that decrease the fidelity and intelligibility of the desired speaker. Reverberant speech can be separated into two components, viz. early speech and late reverberant speech. Recently, various algorithms have been developed to suppress late reverberant speech. One of the main challenges is to develop an estimator for the so-called late reverberant spectral variance (LRSV) which is required by most of these algorithms. In this letter a statistical reverberation model is proposed that takes the energy contribution of the direct-path into account. This model is then used to derive a more general LRSV estimator, which in a particular case reduces to an existing LRSV estimator. Experimental results show that the developed estimator is advantageous in case the source-microphone distance is smaller than the critical distance.
Emanuël A. P. Habets, Sharon Gannot, Israel Cohen
IEEE Signal Process. Lett.3
2009 Multichannel Eigenspace Beamforming in a Reverberant Noisy Environment With Multiple Interfering Speech Signals
abstract
In many practical environments we wish to extract several desired speech signals, which are contaminated by nonstationary and stationary interfering signals. The desired signals may also be subject to distortion imposed by the acoustic room impulse responses (RIRs). In this paper, a linearly constrained minimum variance (LCMV) beamformer is designed for extracting the desired signals from multimicrophone measurements. The beamformer satisfies two sets of linear constraints. One set is dedicated to maintaining the desired signals, while the other set is chosen to mitigate both the stationary and nonstationary interferences. Unlike classical beamformers, which approximate the RIRs as delay-only filters, we take into account the entire RIR [or its respective acoustic transfer function (ATF)]. The LCMV beamformer is then reformulated in a generalized sidelobe canceler (GSC) structure, consisting of a fixed beamformer (FBF), blocking matrix (BM), and adaptive noise canceler (ANC). It is shown that for spatially white noise field, the beamformer reduces to a FBF, satisfying the constraint sets, without power minimization. It is shown that the application of the adaptive ANC contributes to interference reduction, but only when the constraint sets are not completely satisfied. We show that relative transfer functions (RTFs), which relate the desired speech sources and the microphones, and a basis for the interference subspace suffice for constructing the beamformer. The RTFs are estimated by applying the generalized eigenvalue decomposition (GEVD) procedure to the power spectral density (PSD) matrices of the received signals and the stationary noise. A basis for the interference subspace is estimated by collecting eigenvectors, calculated in segments where nonstationary interfering sources are active and the desired sources are inactive. The rank of the basis is then reduced by the application of the orthogonal triangular decomposition (QRD). This procedure relaxes the common requirement for nonoverlapping activity periods of the interference sources. A comprehensive experimental study in both simulated and real environments demonstrates the performance of the proposed beamformer.
Shmulik Markovich-Golan, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.3
2009 Relative Transfer Function Identification Using Convolutive Transfer Function Approximation
abstract
In this paper, we present a relative transfer function (RTF) identification method for speech sources in reverberant environments. The proposed method is based on the convolutive transfer function (CTF) approximation, which enables to represent a linear convolution in the time domain as a linear convolution in the short-time Fourier transform (STFT) domain. Unlike the restrictive and commonly used multiplicative transfer function (MTF) approximation, which becomes more accurate when the length of a time frame increases relative to the length of the impulse response, the CTF approximation enables representation of long impulse responses using short time frames. We develop an unbiased RTF estimator that exploits the nonstationarity and presence probability of the speech signal and derive an analytic expression for the estimator variance. Experimental results show that the proposed method is advantageous compared to common RTF identification methods in various acoustic environments, especially when identifying long RTFs typical to real rooms.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.2
2009 Convolutive Transfer Function Generalized Sidelobe Canceler
abstract
In this paper, we propose a convolutive transfer function generalized sidelobe canceler (CTF-GSC), which is an adaptive beamformer designed for multichannel speech enhancement in reverberant environments. Using a complete system representation in the short-time Fourier transform (STFT) domain, we formulate a constrained minimization problem of total output noise power subject to the constraint that the signal component of the output is the desired signal, up to some prespecified filter. Then, we employ the general sidelobe canceler (GSC) structure to transform the problem into an equivalent unconstrained form by decoupling the constraint and the minimization. The CTF-GSC is obtained by applying a convolutive transfer function (CTF) approximation on the GSC scheme, which is a more accurate and a less restrictive than a multiplicative transfer function (MTF) approximation. Experimental results demonstrate that the proposed beamformer outperforms the transfer function GSC (TF-GSC) in reverberant environments and achieves both improved noise reduction and reduced speech distortion.
Ronen Talmon, Israel Cohen, Sharon Gannot
IEEE Trans. Speech Audio Process.2
2009 Multichannel Seismic Deconvolution Using Markov-Bernoulli Random-Field Modeling
abstract
In this paper, we present an algorithm for multichannel blind deconvolution of seismic signals, which exploits lateral continuity of Earth layers based on Markov-Bernoulli random-field modeling. The reflectivity model accounts for layer discontinuities resulting from splitting, merging, starting, or terminating layers within the region of interest. We define a set of reflectivity states and legal transitions between the reflector configurations of adjacent traces and subsequently apply the Viterbi algorithm for finding the most likely sequences of reflectors that are connected across the traces by legal transitions. The improved performance of the proposed algorithm and its robustness to noise, compared with a competitive algorithm, are demonstrated using simulated and real seismic data examples, in blind and nonblind scenarios.
Alon Heimer, Israel Cohen
IEEE Trans. Geosci. Remote. Sens.2
2008 Dual-microphone speech dereverberation using GARCH modeling
abstract
In this paper, we develop a dual-microphone speech dereverberation algorithm for noisy environments, which is aimed at suppressing late reverberation and background noise. The spectral variance of the late reverberation is obtained with adaptively-estimated direct path compensation. A Markov-switching generalized autoregressive conditional heteroscedasticity (GARCH) model is used to estimate the spectral variance of the desired signal, which includes the direct sound and early reverberation. Experimental results demonstrate the advantage of the proposed algorithm compared to a decision-directed-based algorithm.
Ari Abramson, Emanuël A. P. Habets, Sharon Gannot, Israel Cohen
ICASSP4
2008 Identification of linear systems with adaptive control of the cross-multiplicative transfer function approximation
abstract
In this paper, we extend the cross-multiplicative transfer function (CMTF) approach for improved system identification in the short- time Fourier transform (STFT) domain. The proposed algorithm adaptively controls the number of cross-terms in the CMTF approximation to achieve the minimum mean-square error (mmse) at each iteration. A small number of cross-terms is initially used to achieve fast convergence, and as the adaptation process proceeds, the algorithm gradually increases this number to enhance the steady-state performance. When compared to the conventional multiplicative transfer function (MTF) approach, the resulting algorithm achieves a substantial improvement in steady-state performance, without compromising for slower convergence. Experimental results validate the theoretical derivations and demonstrate the advantage of the proposed approach to acoustic echo cancellation.
Yekutiel Avargel, Israel Cohen
ICASSP2
2008 Multichannel blind seismic deconvolution using dynamic programming
Alon Heimer, Israel Cohen
Signal Process.2
2008 Single-Sensor Audio Source Separation Using Classification and Estimation Approach and GARCH Modeling
abstract
In this paper, we propose a new algorithm for single-sensor audio source separation of speech and music signals, which is based on generalized autoregressive conditional heteroscedasticity (GARCH) modeling of the speech signals and Gaussian mixture modeling (GMM) of the music signals. The separation of the speech from the music signal is obtained by a simultaneous classification and estimation approach, which enables one to control the tradeoff between residual interference and signal distortion. Experimental results on mixtures of speech and piano music signals have yielded an improved source separation performance compared to using Gaussian mixture models for both signals. The tradeoff between signal distortion and residual interference is controlled by adjusting some cost parameters, which are shown to determine the missed and false detection rates in the proposed classification and estimation approach.
Ari Abramson, Israel Cohen
IEEE Trans. Speech Audio Process.2
2008 Adaptive System Identification in the Short-Time Fourier Transform Domain Using Cross-Multiplicative Transfer Function Approximation
abstract
In this paper, we introduce cross-multiplicative transfer function (CMTF) approximation for modeling linear systems in the short-time Fourier transform (STFT) domain. We assume that the transfer function can be represented by cross-multiplicative terms between distinct subbands. We investigate the influence of cross-terms on a system identifier implemented in the STFT domain and derive analytical relations between the noise level, data length, and number of cross-multiplicative terms, which are useful for system identification. As more data becomes available or as the noise level decreases, additional cross-terms should be considered and estimated to attain the minimal mean-square error (mse). A substantial improvement in performance is then achieved over the conventional multiplicative transfer function (MTF) approximation. Furthermore, we derive explicit expressions for the transient and steady-state mse performances obtained by adaptively estimating the cross-terms. As more cross-terms are estimated, a lower steady-state mse is achieved, but the algorithm then suffers from slower convergence. Experimental results validate the theoretical derivations and demonstrate the effectiveness of the proposed approach as applied to acoustic echo cancellation.
Yekutiel Avargel, Israel Cohen
IEEE Trans. Speech Audio Process.2
2008 Joint Dereverberation and Residual Echo Suppression of Speech Signals in Noisy Environments
abstract
Hands-free devices are often used in a noisy and reverberant environment. Therefore, the received microphone signal does not only contain the desired near-end speech signal but also interferences such as room reverberation that is caused by the near-end source, background noise and a far-end echo signal that results from the acoustic coupling between the loudspeaker and the microphone. These interferences degrade the fidelity and intelligibility of near-end speech. In the last two decades, postfilters have been developed that can be used in conjunction with a single microphone acoustic echo canceller to enhance the near-end speech. In previous works, spectral enhancement techniques have been used to suppress residual echo and background noise for single microphone acoustic echo cancellers. However, dereverberation of the near-end speech was not addressed in this context. Recently, practically feasible spectral enhancement techniques to suppress reverberation have emerged. In this paper, we derive a novel spectral variance estimator for the late reverberation of the near-end speech. Residual echo will be present at the output of the acoustic echo canceller when the acoustic echo path cannot be completely modeled by the adaptive filter. A spectral variance estimator for the so-called late residual echo that results from the deficient length of the adaptive filter is derived. Both estimators are based on a statistical reverberation model. The model parameters depend on the reverberation time of the room, which can be obtained using the estimated acoustic echo path. A novel postfilter is developed which suppresses late reverberation of the near-end speech, residual echo and background noise, and maintains a constant residual background noise level. Experimental results demonstrate the beneficial use of the developed system for reducing reverberation, residual echo, and background noise.
Emanuël A. P. Habets, Sharon Gannot, Israel Cohen, P. Sommen
IEEE Trans. Speech Audio Process.3
2008 Dual-Source Transfer-Function Generalized Sidelobe Canceller
abstract
Full-duplex hands-free man/machine interface often suffers from directional nonstationary interference, such as a competing speaker, as well as stationary interferences which may comprise both directional and nondirectional signals. The transfer-function generalized sidelobe canceller (TF-GSC) exploits the nonstationarity of the speech signal to enhance it when the undesired interfering signals are stationary. Unfortunately, the assumptions leading to the derivation of the TF-GSC are violated when a nonstationary interference is present. In this paper, we propose an adaptive beamformer, based on the TF-GSC, that is suitable for cancelling nonstationary interferences in noisy reverberant environments. We modify two of the TF-GSC components to enable suppression of the nonstationary undesired signal. A modified fixed beamformer (FBF) is designed to block the nonstationary interfering signal while maintaining the desired speech signal. A modified blocking matrix (BM) is designed to block both the desired signal and the nonstationary interference. We introduce a novel method for updating the blocking matrix in double talk scenarios, which exploits the nonstationarity of both the desired and interfering speech signals. Experimental results demonstrate the performance of the proposed algorithm in noisy and reverberant environments and show its superiority over the original TF-GSC.
Gal Reuven, Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.3
2007 Enhancement of Speech Signals Under Multiple Hypotheses using an Indicator for Transient Noise Presence
abstract
In this paper, we formulate a speech enhancement problem under multiple hypotheses, assuming an indicator or detector for the transient noise presence is available in the short-time Fourier transform (STFT) domain. Hypothetical presence of speech or transient noise is considered in the observed spectral coefficients, and cost parameters control the trade-off between speech distortion and residual transient noise. An optimal estimator, which minimizes the mean-square error of the log-spectral amplitude, is derived, while taking into account the probability of erroneous detection. Experimental results demonstrate the improved performance in transient noise suppression, compared to using the optimally-modified log-spectral amplitude estimator.
Ari Abramson, Israel Cohen
ICASSP (4)2
2007 Multichannel Acoustic Echo Cancellation and Noise Reduction in Reverberant Environments using the Transfer-Function GSC
abstract
In this paper, we present a multi-channel acoustic echo canceller that is integrated into the transfer-function generalized sidelobe canceller (TF-GSC). The proposed scheme consists of a primary TF-GSC, for dealing with the noise interferences, and a secondary modified TF-GSC, for dealing with the echo cancellation. The secondary TF-GSC includes an echo canceller embedded within a replica of the primary TF-GSC components. Experimental results demonstrate improved performance compared to cascade schemes of acoustic echo cancellation and adaptive beamforming.
Gal Reuven, Sharon Gannot, Israel Cohen
ICASSP (1)3
2007 Detection of anomalies in texture images using multi-resolution random field models
Lior Shadhan, Israel Cohen
Signal Process.2
2007 Special issue on Speech Enhancement
Philipos C. Loizou, Israel Cohen, Sharon Gannot, Kuldip K. Paliwal
Speech Commun.2
2007 Performance analysis of dual source transfer-function generalized sidelobe canceller
Gal Reuven, Sharon Gannot, Israel Cohen
Speech Commun.3
2007 Joint noise reduction and acoustic echo cancellation using the transfer-function generalized sidelobe canceller
Gal Reuven, Sharon Gannot, Israel Cohen
Speech Commun.3
2007 On Multiplicative Transfer Function Approximation in the Short-Time Fourier Transform Domain
abstract
The multiplicative transfer function (MTF) approximation is widely used for modeling a linear time invariant system in the short-time Fourier transform (STFT) domain. It relies on the assumption of a long analysis window compared with the length of the system impulse response. In this paper, we investigate the influence of the analysis window length on the performance of a system identifier that utilizes the MTF approximation. We derive analytic expressions for the minimum mean-square error (MMSE) in the STFT domain and show that the system identification performance does not necessarily improve by increasing the length of the analysis window. The optimal window length, that achieves the MMSE, depends on the signal-to-noise ratio and the length of the input signal. The theoretical analysis is supported by simulation results
Yekutiel Avargel, Israel Cohen
IEEE Signal Process. Lett.2
2007 Audio Packet Loss Concealment in a Combined MDCT-MDST Domain
abstract
Audio streaming applications have become very popular in recent years, owing to their low cost and convenience. However, during network congestions, data packets are often delayed or discarded, creating an annoying gap in the streamed media. This letter presents a new approach to audio packet loss concealment designed for MPEG-audio streaming applications. In a previous work, we introduced a receiver-based concealment algorithm based on applying the gapped-data amplitude and phase estimation (GAPES) interpolation algorithm in the discrete short-time Fourier transform (DSTFT) complex domain and obtained better results compared to past methods. The current approach applies the same algorithm on a different complex domain, formed from combining the modified discrete cosine transform (MDCT) domain as its real part and the modified discrete sine transform (MDST) domain as the imaginary part. The new approach significantly reduces the complexity demands while maintaining similar high-quality results.
Hadas Ofir, David Malah, Israel Cohen
IEEE Signal Process. Lett.3
2007 Simultaneous Detection and Estimation Approach for Speech Enhancement
abstract
In this paper, we present a simultaneous detection and estimation approach for speech enhancement. A detector for speech presence in the short-time Fourier transform domain is combined with an estimator, which jointly minimizes a cost function that takes into account both detection and estimation errors. Cost parameters control the tradeoff between speech distortion, caused by missed detection of speech components and residual musical noise resulting from false-detection. Furthermore, a modified decision-directed a priori signal-to-noise ratio (SNR) estimation is proposed for transient-noise environments. Experimental results demonstrate the advantage of using the proposed simultaneous detection and estimation approach with the proposed a priori SNR estimator, which facilitate suppression of transient noise with a controlled level of speech distortion.
Ari Abramson, Israel Cohen
IEEE Trans. Speech Audio Process.2
2007 System Identification in the Short-Time Fourier Transform Domain With Crossband Filtering
abstract
In this paper, we investigate the influence of crossband filters on a system identifier implemented in the short-time Fourier transform (STFT) domain. We derive analytical relations between the number of crossband filters, which are useful for system identification in the STFT domain, and the power and length of the input signal. We show that increasing the number of crossband filters not necessarily implies a lower steady-state mean-square error (mse) in subbands. The number of useful crossband filters depends on the power ratio between the input signal and the additive noise signal. Furthermore, it depends on the effective length of input signal employed for system identification, which is restricted to enable tracking capability of the algorithm during time variations in the system. As the power of input signal increases or as the time variations in the system become slower, a larger number of crossband filters may be utilized. The proposed subband approach is compared to the conventional fullband approach and to the commonly used subband approach that relies on multiplicative transfer function (MTF) approximation. The comparison is carried out in terms of mse performance and computational complexity. Experimental results verify the theoretical derivations and demonstrate the relations between the number of useful crossband filters and the power and length of the input signal
Yekutiel Avargel, Israel Cohen
IEEE Trans. Speech Audio Process.2
2007 Anomaly Detection Based on Wavelet Domain GARCH Random Field Modeling
abstract
One-dimensional Generalized Autoregressive Conditional Heteroscedasticity (GARCH) model is widely used for modeling financial time series. Extending the GARCH model to multiple dimensions yields a novel clutter model which is capable of taking into account important characteristics of a wavelet-based multiscale feature space, namely heavy-tailed distributions and innovations clustering as well as spatial and scale correlations. We show that the multidimensional GARCH model generalizes the casual Gauss Markov random field (GMRF) model, and we develop a multiscale matched subspace detector (MSD) for detecting anomalies in GARCH clutter. Experimental results demonstrate that by using a multiscale MSD under GARCH clutter modeling, rather than GMRF clutter modeling, a reduced false-alarm rate can be achieved without compromising the detection rate
Amir Noiboar, Israel Cohen
IEEE Trans. Geosci. Remote. Sens.2
2006 Asymptotic Stationarity of Markov-Switching Time-Frequency Garch Processes
abstract
Conditions for asymptotic wide-sense stationarity of generalized autoregressive conditional heteroscedasticity (GARCH) processes with regime-switching are necessary for ensuring finite second moments. In this paper, we introduce a stationarity analysis for the Markov-switching time-frequency GARCH (MSTF-GARCH) model which has been recently introduced for modeling nonstationary signals in the time-frequency domain. We obtain a recursive vector form for the unconditional variance by using a representative matrix which is constructed from both the GARCH parameters of each regime, and the regimes' transition probabilities. We show that constraining the spectral radius of that matrix to be less than one is both necessary and sufficient for asymptotic wide-sense stationarity. The generated matrix is also shown to be useful for deriving the asymptotic covariance matrix of the process
Ari Abramson, Israel Cohen
ICASSP (3)2
2006 Speech spectral modeling and enhancement based on autoregressive conditional heteroscedasticity models
Israel Cohen
Signal Process.1
2006 State smoothing in Markov-switching time-frequency GARCH models
abstract
In this letter, we propose a state smoothing algorithm for path-dependent Markov-switching generalized autoregressive conditional heteroscedasticity (GARCH) processes. Our smoothing technique extends the forward-backward recursions of Chang and Hancock and the stable backward recursion of Lindgren, Askar and Derin. We derive two recursive steps for the evaluation of conditional densities of future observations. The first step is an upward recursion that manipulates the future observations for the evaluation of their conditional densities, and the second step is a backward recursion that integrates over the possible future paths. Experimental results demonstrate the improvement in performance, compared to using causal estimation.
Ari Abramson, Israel Cohen
IEEE Signal Process. Lett.2
2005 Supergaussian GARCH models for speech signals
abstract
In this paper, we introduce supergaussian generalized autoregres- sive conditional heteroscedasticity (GARCH) models for speech signals in the short-time Fourier transform (STFT) domain. We address the problem of speech enhancement, and show that esti- mating the variances of the STFT expansion coefficients based on GARCH models yields higher speech quality than by using the decision-directed method, whether the fidelity criterion is mini- mum mean-squared error (MMSE) of the spectral coefficients or MMSE of the log-spectral amplitude (LSA). Furthermore, while a Gaussian model is inferior to Gamma and Laplacian models when estimating the variances by the decision-directed method, a Gaus- sian model is superior when using the GARCH modeling method. This facilitates MMSE-LSA estimation, while taking into consid- eration the heavy-tailed distribution. tudes tend to follow large magnitudes and small magnitudes tend to follow small magnitudes, while the phase is unpredictable. This paper summarizes the main results of (8). We present supergaussian GARCH models for speech signals in the STFT domain. We address the problem of spectral enhancement of noisy speech, and consider eight different speech enhancement algorithms, as summarized in Table 1. The statistical model is either Gaussian, Gamma or Laplacian; the spectral variance is estimated based on either the proposed GARCH models or the decision-directed method of Ephraim and Malah (2); the fidelity criteria include MMSE of the STFT coefficients and MMSE of the LSA. We show that estimating the variance by the GARCH model- ing method yields lower log-spectral distortion (LSD) and higher Perceptual Evaluation of Speech Quality (PESQ) scores (ITU-T P.862) than by using the decision-directed method. Furthermore, while a Gaussian model is inferior to Gamma and Laplacian mod- els if the speech variance is estimated by the decision-directed method, a Gaussian model is superior in the case speech variance is estimated by using the GARCH modeling method. This facili- tates MMSE-LSA estimation, while taking into consideration the heavy-tailed distribution. Speech spectrograms and informal lis- tening tests confirm that the quality of the enhanced speech ob- tained by using the GARCH modeling method is better than that obtainable by using the decision-directed method. In Sec. 2, we introduce the statistical models. In Sec. 3, we address the speech enhancement problem. In Sec. 4, we derive estimators for the spectral variances. Finally, in Sec. 5, we evalu- ate the performances of MMSE and MMSE-LSA estimators under Gaussian, Gamma and Laplacian models.
Israel Cohen
INTERSPEECH1
2005 Anomaly subspace detection based on a multi-scale Markov random field model
Arnon Goldman, Israel Cohen
Signal Process.2
2005 Speech enhancement using super-Gaussian speech models and noncausal a priori SNR estimation
Israel Cohen
Speech Commun.1
2005 Relaxed Statistical Model for Speech Enhancement and a Priori SNR Estimation
abstract
In this paper, we propose a statistical model for speech enhancement that takes into account the time-correlation between successive speech spectral components. It retains the simplicity associated with the Gaussian statistical model, and enables the extension of existing algorithms to noncausal estimation. The sequence of speech spectral variances is a random process, which is generally correlated with the sequence of speech spectral magnitudes. Causal and noncausal estimators for the a priori SNR are derived in agreement with the model assumptions and the estimation of the speech spectral components. We show that a special case of the causal estimator degenerates to a "decision-directed" estimator with a time-varying frequency-dependent weighting factor. Experimental results demonstrate the improved performance of the proposed algorithms.
Israel Cohen
IEEE Trans. Speech Audio Process.1
2004 On the decision-directed estimation approach of Ephraim and Malah
abstract
The decision-directed approach of Y. Ephraim and D. Malah (IEEE Trans. Acoustics, Speech and Signal Proc., vol.ASSP-32, no.6, p.1109-21, 1984) is widely used for a priori SNR estimation and speech enhancement. However, it conflicts with common model assumptions. We propose recursive estimators for the a priori SNR and the speech spectral components. We introduce a novel statistical model that takes into account the time-correlation between successive speech spectral components, while keeping the resulting algorithms simple. This model provides new insight into the decision-directed approach, and enables the extension of existing speech enhancement algorithms to noncausal estimation. The causal a priori SNR estimator degenerates, as a special case, to a "decision-directed" estimator with a time-varying frequency-dependent weighting factor. The noncausal estimator is capable of discriminating between speech onsets and noise irregularities, achieving lower levels of both musical noise and speech distortion.
Israel Cohen
ICASSP (1)1
2004 Modeling speech signals in the time-frequency domain using GARCH
Israel Cohen
Signal Process.1
2004 Anomaly detection based on an iterative local statistics approach
Arnon Goldman, Israel Cohen
Signal Process.2
2004 Identification of speech source coupling between sensors in reverberant noisy environments
abstract
An important component of a multichannel hands-free communication system is the identification of the coupling between sensors in response to a desired speech signal. In this letter, a system identification approach adapted to speech signals is proposed. A weighted least-squares optimization criterion is introduced, which incorporates an indicator function for the presence of the desired signal in the observed signals. We show that compared to a competing nonstationarity-based method, a significantly smaller error variance is achievable.
Israel Cohen
IEEE Signal Process. Lett.1
2004 Speech enhancement using a noncausal a priori SNR estimator
abstract
We propose a noncausal estimator for the a priori signal-to-noise ratio (SNR), and a corresponding noncausal speech enhancement algorithm. In contrast to the decision-directed estimator of Ephraim and Malah (1984), the noncausal estimator is capable of discriminating between speech onsets and noise irregularities. Onsets of speech are better preserved, while a further reduction of musical noise is achieved. Experimental results show that the noncausal estimator yields a higher improvement in the segmental SNR, lower log-spectral distortion, and better Perceptual Evaluation of Speech Quality scores (PESQ, ITU-T P.862).
Israel Cohen
IEEE Signal Process. Lett.1
2004 Relative transfer function identification using speech signals
abstract
An important component of a multichannel hands-free communication system is the identification of the relative transfer function between sensors in response to a desired source signal. In this paper, a robust system identification approach adapted to speech signals is proposed. A weighted least-squares optimization criterion is introduced, which considers the uncertainty of the desired signal presence in the observed signals. An asymptotically unbiased estimate for the system's transfer function is derived, and a corresponding recursive online implementation is presented. We show that compared to a competing nonstationarity-based method, a smaller error variance is achieved and generally shorter observation intervals are required. Furthermore, in the case of a time-varying system, faster convergence and higher reliability of the system identification are obtained by using the proposed method than by using the nonstationarity-based method. Evaluation of the proposed system identification approach is performed under various noise conditions, including simulated stationary and nonstationary white Gaussian noise, and car interior noise in real pseudo-stationary and nonstationary environments. The experimental results confirm the advantages of proposed approach.
Israel Cohen
IEEE Trans. Speech Audio Process.1
2004 Speech enhancement based on the general transfer function GSC and postfiltering
abstract
In speech enhancement applications microphone array postfiltering allows additional reduction of noise components at a beamformer output. Among microphone array structures the recently proposed general transfer function generalized sidelobe canceller (TF-GSC) has shown impressive noise reduction abilities in a directional noise field, while still maintaining low speech distortion. However, in a diffused noise field less significant noise reduction is obtainable. The performance is even further degraded when the noise signal is nonstationary. In this contribution we propose three postfiltering methods for improving the performance of microphone arrays. Two of which are based on single-channel speech enhancers and making use of recently proposed algorithms concatenated to the beamformer output. The third is a multichannel speech enhancer which exploits noise-only components constructed within the TF-GSC structure. This work concentrates on the assessment of the proposed postfiltering structures. An extensive experimental study, which consists of both objective and subjective evaluation in various noise fields, demonstrates the advantage of the multichannel postfiltering compared to the single-channel techniques.
Sharon Gannot, Israel Cohen
IEEE Trans. Speech Audio Process.2
2003 Two-channel signal detection and speech enhancement based on the transient beam-to-reference ratio
abstract
In reverberant and noisy environments, multichannel systems are designed for spatially filtering interfering signals coming from undesired directions. In case of incoherent or diffuse noise fields, beamforming alone does not provide sufficient noise reduction, and post-filtering is normally required. In this paper, we present a two-channel post-filtering approach for signal detection and speech enhancement. A mild assumption is made, that a desired signal component is stronger at the beamformer output than at the reference noise signal, and a noise component is stronger at the reference signal. The ratio between the transient power at the beamformer output and the transient power at the reference noise signal is used for indicating whether such a transient is desired or interfering. Experimental results demonstrate the usefulness of the proposed approach in a car environment.
Israel Cohen, Baruch Berdugo
ICASSP (5)1
2003 Speech enhancement based on the general transfer function GSC and postfiltering
abstract
In speech enhancement applications, microphone array postfiltering allows additional reduction of noise components at a beamformer output. Among microphone array structures, the recently proposed general transfer function generalized sidelobe canceller (TF-GSC) has shown impressive noise reduction abilities in a directional noise field, while still maintaining low speech distortion. However, in a diffused noise field, less significant noise reduction is obtainable. The performance is even further degraded when the noise is nonstationary. We present three postfiltering methods for improving the performance of microphone arrays. Two of them are based on single-channel speech enhancers and make use of recently proposed algorithms concatenated to the beamformer output. The third is a multichannel speech enhancer which exploits noise-only components constructed within the TF-GSC structure. An experimental study, which consists of both objective and subjective evaluation in various noise fields, demonstrates the advantage of the multi-channel postfiltering compared to single-channel techniques.
Sharon Gannot, Israel Cohen
ICASSP (1)2
2003 Multichannel signal detection based on the transient beam-to-reference ratio
abstract
We present a multichannel signal detection approach that is particularly advantageous in nonstationary noise environments. A beamformer is realistically assumed to have a steering error, a blocking matrix that is unable to block all of the desired signal components, and a noise canceler that is adapted to the pseudostationary noise, but not modified during transient interference. Signal components are detected at the beamformer output based on a measure of their local nonstationarity, and discriminated from transient noise components based on the transient beam-to-reference ratio.
Israel Cohen, Baruch Berdugo
IEEE Signal Process. Lett.1
2003 Noise spectrum estimation in adverse environments: improved minima controlled recursive averaging
abstract
Noise spectrum estimation is a fundamental component of speech enhancement and speech recognition systems. We present an improved minima controlled recursive averaging (IMCRA) approach, for noise estimation in adverse environments involving nonstationary noise, weak speech components, and low input signal-to-noise ratio (SNR). The noise estimate is obtained by averaging past spectral power values, using a time-varying frequency-dependent smoothing parameter that is adjusted by the signal presence probability. The speech presence probability is controlled by the minima values of a smoothed periodogram. The proposed procedure comprises two iterations of smoothing and minimum tracking. The first iteration provides a rough voice activity detection in each frequency band. Then, smoothing in the second iteration excludes relatively strong speech components, which makes the minimum tracking during speech activity robust. We show that in nonstationary noise environments and under low SNR conditions, the IMCRA approach is very effective. In particular, compared to a competitive method, it obtains a lower estimation error, and when integrated into a speech enhancement system achieves improved speech quality and lower residual noise.
Israel Cohen
IEEE Trans. Speech Audio Process.1
2003 Analysis of two-channel generalized sidelobe canceller (GSC) with post-filtering
abstract
In this paper, we analyze a two-channel generalized sidelobe canceller with post-filtering in nonstationary noise environments. The post-filtering includes detection of transients at the beamformer output and reference signal, a comparison of their transient power, estimation of the signal presence probability, estimation of the noise spectrum, and spectral enhancement for minimizing the mean-square error of the log-spectra. Transients are detected based on a measure of their local nonstationarity, and classified as desired or interfering based on the transient beam-to-reference ratio. We introduce a transient discrimination quality measure, which quantifies the beamformer's capability to recognize noise transients as distinct from signal transients. Evaluating this measure in various noise fields shows that desired and interfering transients can generally be differentiated within a wide range of frequencies. To further improve the transient noise reduction at low and high frequencies in case the signal is wideband, we estimate for each time frame a global likelihood of signal presence. The global likelihood is associated with the transient beam-to-reference ratios in frequencies, where the transient discrimination quality is high. Experimental results demonstrate the usefulness of the proposed approach in various car environments.
Israel Cohen
IEEE Trans. Speech Audio Process.1
2002 Microphone array post-filtering for non-stationary noise suppression
abstract
Microphone array post-filtering allows additional reduction of noise components at a beamformer output. Existing techniques are either restricted to classical delay-and-sum beamformers, or are based on single-channel speech enhancement algorithms that are inefficient at attenuating highly non-stationary noise components. In this paper, we introduce a microphone array post-filtering approach, applicable to adaptive beamformer, that differentiates non-stationary noise components from speech components. The ratio between the transient power at the beamformer primary output and the transient power at the reference noise signals is used for indicating whether such a transient is desired or interfering. Based on a Gaussian statistical model and combined with an appropriate spectral enhancement technique, a significantly reduced level of non-stationary noise is achieved without further distorting speech components. Experimental results demonstrate the effectiveness of the proposed method.
Israel Cohen, Baruch Berdugo
ICASSP1
2002 Optimal speech enhancement under signal presence uncertainty using log-spectral amplitude estimator
abstract
We present an optimally modified log-spectral amplitude estimator, which minimizes the mean-square error of the log-spectra for speech signals under signal presence uncertainty. We propose an estimator for the a priori signal-to-noise ratio (SNR), and introduce an efficient estimator for the a priori speech absence probability. Speech presence probability is estimated for each frequency bin and each frame by a soft-decision approach, which exploits the strong correlation of speech presence in neighboring frequency bins of consecutive frames. Objective and subjective evaluation confirm superiority in noise suppression and quality of the enhanced speech.
Israel Cohen
IEEE Signal Process. Lett.1
2002 Noise estimation by minima controlled recursive averaging for robust speech enhancement
abstract
In this letter, we introduce a minima controlled recursive averaging (MCRA) approach for noise estimation. The noise estimate is given by averaging past spectral power values and using a smoothing parameter that is adjusted by the signal presence probability in subbands. The presence of speech in subbands is determined by the ratio between the local energy of the noisy speech and its minimum within a specified time window. The noise estimate is computationally efficient, robust with respect to the input signal-to-noise ratio (SNR) and type of underlying additive noise, and characterized by the ability to quickly follow abrupt changes in the noise spectrum.
Israel Cohen, Baruch Berdugo
IEEE Signal Process. Lett.1
2001 On speech enhancement under signal presence uncertainty
abstract
In this paper, we present an optimally-modified log-spectral amplitude estimator, which minimizes the mean-square error of the log-spectra for speech signals under signal presence uncertainty. The spectral gain function is obtained as a weighted geometric mean of the hypothetical gains associated with signal presence and absence. The exponential weight of each hypothetical gain is its corresponding probability, conditioned on the observed signal. We introduce an efficient estimation approach for the a priori signal absence probability in each frequency bin, which exploits the strong correlation of speech presence in neighboring frequency bins of consecutive frames. Objective and subjective evaluation confirm superiority in noise suppression and quality of the enhanced speech.
Israel Cohen
ICASSP1
2001 Enhancement of speech using bark-scaled wavelet packet decomposition
abstract
In this paper, ve propose a speech enhancement system, vhich integrates a bark-scaled vavelet packet decompo- sition (BS-WPD), a soft-decision gain modification and a "magnitude" decision-directed estimation technique. The BS-WPD provides an overcomplete auditory representation, having a higher frequency resolution than the critical band decomposition. Speech is estimated by Wiener filtering in the vavelet packet domain, modified by the signal presence probability. We introduce a "magnitude" decision-directed estimator for the variance of speech, vhich is closely related to the decisiondirected estimator of Ephraim and Malah. This estimator achieves, in the established process, a better tradeoff betveen noise reduction and signal distortion. The proposed enhancement algorithm is tested vith various noise types, and compared to a conventional log-spectral amplitude estimator. We shov that noise can be further suppressed, vhile preserving its natural structure and the intelligibility and quality of the speech components.
Israel Cohen
INTERSPEECH1
2001 Speech enhancement for non-stationary noise environments
Israel Cohen, Baruch Berdugo
Signal Process.1
1999 Translation-invariant denoising using the minimum description length criterion
Israel Cohen, Shalom Raz, David Malah
Signal Process.1
1999 Adaptive suppression of Wigner interference-terms using shift-invariant wavelet packet decompositions
Israel Cohen, Shalom Raz, David Malah
Signal Process.1
1997 Eliminating interference terms in the Wigner distribution using extended libraries of bases
abstract
The Wigner distribution (WD) possesses a number of desirable mathematical properties relevant to time-frequency analysis. However, the presence of interference terms renders the WD of multicomponent signals extremely difficult to interpret. We propose an adaptive decomposition of the WD using extended libraries of orthonormal bases. A prescribed signal is expanded on a basis of adapted waveforms, that best match the signal components, and subsequently transformed into the Wigner domain. The interference terms are controlled by thresholding the cross WD of interactive basis functions according to their degree of adjacency in an idealized time-frequency plane. This measure is implicitly adapted to the local distribution of the signal, thus compensating for a global nonadaptive threshold. In particular we focus on a shift-invariant decomposition in an extended library of wavelet packets. The resulting modified distribution achieves high time-frequency resolution, and is superior in eliminating interference terms associated with bilinear distributions.
Israel Cohen, Shalom Raz, David Malah
ICASSP1
1997 Orthonormal shift-invariant adaptive local trigonometric decomposition
Israel Cohen, Shalom Raz, David Malah
Signal Process.1
1997 Orthonormal shift-invariant wavelet packet decomposition and representation
Israel Cohen, Shalom Raz, David Malah
Signal Process.1
1995 Shift invariant wavelet packet bases
abstract
A shifted wavelet packet (SWP) library, containing all the time shifted wavelet packet bases, is defined. A corresponding shift-invariant wavelet packet decomposition (SIWPD) search algorithm for a "best basis" is introduced. The search algorithm is representable by a binary tree, in which a node symbolizes an appropriate subspace of the original signal. We prove that the resultant "best basis" is orthonormal and the associated expansion, characterized by the lowest "information cost", is shift-invariant. The shift-invariance stems from an additional degree of freedom, generated at the decomposition stage, and incorporated into the search algorithm. We prove that for any subspace it suffices to consider one of two alternative decompositions, made feasible by the SWP library. The computational complexity of SIWPD may be controlled at the expense of the attained information cost, to an extent of O(2Nlog/sub 2/N).
Israel Cohen, Shalom Raz, David Malah
ICASSP1
1995 Shift-invariant adaptive local trigonometric decomposition
abstract
A general formulation of shift-invariant "best-basis" expansions is presented. Specifically, we construct an extended library of smooth local trigonometric bases, and introduce a suitable "best-basis" search algorithm. We prove that the resultant decomposition is shift-invariant, orthonormal and characterized by a reduced information cost. The shift-invariance is derived from an adaptive relative shift of expansions in distinct resolution levels. We show that at any resolution level it suffices to examine and select one of two relative shift options a zero shift or a 2-e-1 shift. A variable folding operator, whose polarity is locally adapted to the parity properties of the signal, extra enhances the representation.
Israel Cohen, Shalom Raz, David Malah
EUROSPEECH1