Baruch Berdugo

dblp:78/6380 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
YearPublicationVenuePosition
2024 A User-Centric Approach for Deep Residual-Echo Suppression in Double-Talk
abstract
We introduce a user-centric residual-echo suppression (URES) framework in double-talk. This framework receives a user operating point (UOP) that consists of two metric values: the residual echo suppression level (RESL) and the desired speech-maintained level (DSML) that the user expects from the RES outcome. Then, the URES pipeline undergoes three stages. Firstly, we consider a deep RES model with a tunable design parameter that balances between the RESL and DSML and utilizes 101 pre-trained instances of this model, each with a different design parameter value. Thus, an identical input is expected to generate a different pair of RESL and DSML values in the prediction of every instance. Second, every prediction is separately fed to a subsequent pre-trained deep model instance that estimates the RESL and DSML of the prediction since these metrics depend on unavailable information in practice. Lastly, each pair of RESL and DSML estimates is compared with the UOP. The pairs that match the UOP up to a given tolerance threshold are narrowed down to the prediction with the maximal acoustic-echo cancellation mean-opinion score (AECMOS), which is the output of the URES system. This suggested framework holds three prominent advantages introduced in this study: it generates an RES output with RESL and DSML that match a UOP, supports near-real-time tracking of UOP changes, and applies AECMOS maximization. Experimental results consider 60 h of varied real and synthetic data. Average results can achieve an AECMOS subjectively considered excellent with RESL and DSML deviations of roughly 2 dB from the UOP. Any UOP adjustment can be tracked in less than 40 ms with a real-time factor of 1.92, but due to the high computational resources demanded by the framework, this is enabled on-edge only with high-end dedicated hardware, which limits general availability.
Amir Ivry, Israel Cohen, Baruch Berdugo
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Deep Adaptation Control for Acoustic Echo Cancellation
abstract
We propose a general framework for adaptation control using deep neural networks (NNs) and apply it to acoustic echo cancellation (AEC). First, the optimal step-size that controls the adaptation is derived offline by solving a constrained nonlinear optimization problem that minimizes the adaptive filter misadjustment. Then, a deep NN is trained to learn the relation between the input data and the optimal step-size. In real-time, the NN infers the optimal step-size from streaming data and feeds it to an NLMS filter for AEC. This data-driven method makes no assumptions on the acoustic setup and is entirely non-parametric. Experiments with 100 h of real and synthetic data show that the proposed method outperforms the competition in echo cancellation, speech distortion, and convergence during both single-talk and double-talk.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP3
2022 Off-the-Shelf Deep Integration For Residual-Echo Suppression
abstract
Residual-echo suppression (RES) systems suppress the echo and preserve the speech from a mixture of the two. In hands-free speech communication, RES may also be addressed as a source separation (SS) or speech enhancement (SE) problem, where the echo can be manipulated as an interfering speech signal. In this study, we fine-tune three pre-trained deep learning-based systems originally designed for RES, SS, and SE, and show that the best performing system for the task of RES varies with respect to the acoustic conditions. Then, we propose a real-time data-driven integration of these systems, where a neural network continuously tracks the system that achieves the best performance during both single-talk and double-talk periods. Experiments with 100 h of real and synthetic data show that the integrated system outperforms each individual system in terms of echo suppression and speech distortion in various acoustic environments.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP3
2022 Objective Metrics to Evaluate Residual-Echo Suppression During Double-Talk in the Stereophonic Case
Amir Ivry, Israel Cohen, Baruch Berdugo
INTERSPEECH3
2022 Constant-Beamwidth Beamforming With Nonuniform Concentric Ring Arrays
abstract
Solutions for frequency-invariant beamforming with concentric circular arrays are used in many applications, including microphone arrays and audio communication. However, existing methods focus on the azimuth-beamwidth and consider uniformly spaced circular arrays. This paper presents an approach to designing a constant elevation-beamwidth beamformer for nonuniform concentric ring arrays. The suggested beamformer aims to maximize the array directivity factor while maintaining fixed beamwidth over a wide range of frequencies. The introduced methodology simultaneously selects the ring placements and designs the beamformer weights that achieve optimal performance. We present the problem constraints and cost function, resulting in a sparse beamformer design. In addition, a uniformly spaced beamformer is incorporated into the optimization cost function to attain improved performance. Time-domain implementation of the ideal beamformer is also presented in this work. Furthermore, in the proposed array configuration, all sensors on each ring share the same weight value. Thus, the computational complexity and the physical hardware required in a physical setup are significantly reduced. Experimental results demonstrate the advantages of the nonuniform beamformer compared to a uniform beamformer in terms of directivity index, white noise gain, and sidelobe attenuation.
Avital Kleiman, Israel Cohen, Baruch Berdugo
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Deep Residual Echo Suppression With A Tunable Tradeoff Between Signal Distortion And Echo Suppression
abstract
In this paper, we propose a residual echo suppression method using a UNet neural network that directly maps the outputs of a linear acoustic echo canceler to the desired signal in the spectral domain. This system embeds a design parameter that allows a tunable tradeoff between the desired-signal distortion and residual echo suppression in double-talk scenarios. The system employs 136 thousand parameters, and requires 1.6 Giga floating-point operations per second and 10 Mega-bytes of memory. The implementation satisfies both the timing requirements of the AEC challenge and the computational and memory limitations of on-device applications. Experiments are conducted with 161 h of data from the AEC challenge database and from real independent recordings. We demonstrate the performance of the proposed system in real-life conditions and compare it with two competing methods regarding echo suppression and desired-signal distortion, generalization to various environments, and robustness to high echo levels.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP3
2021 Window Beamformer for Sparse Concentric Circular Array
abstract
This work proposes a practical fixed beamformer for a sparse concentric circular array (CCA) of microphones. Firstly, a method is proposed which enables the reduction of the number of microphones in a CCA based on the optimal placement of the microphones to maximize the directivity-factor (DF). Thereafter, a beamformer is derived which strives to obtain a beamwidth that is not smaller than a desired target beamwidth (in both the elevation and azimuth) without neglecting the DF and white-noise-gain (WNG) across the speech spectrum. The proposed beamformer weights the microphones based on their corresponding ring number or radius, and further uses a Gaussian-window function to emphasize the microphones which are aligned with the direction-of-arrival of the signal. The results showcase the superiority of the proposed Gaussian-window (GW) beamformer over known solutions.
Rajib Sharma, Israel Cohen, Baruch Berdugo
ICASSP3
2021 Nonlinear Acoustic Echo Cancellation with Deep Learning
abstract
We propose a nonlinear acoustic echo cancellation system, which aims to model the echo path from the far-end signal to the near-end microphone in two parts. Inspired by the physical behavior of modern hands-free devices, we first introduce a novel neural network architecture that is specifically designed to model the nonlinear distortions these devices induce between receiving and playing the far-end signal. To account for variations between devices, we construct this network with trainable memory length and nonlinear activation functions that are not parameterized in advance, but are rather optimized during the training stage using the training data. Second, the network is succeeded by a standard adaptive linear filter that constantly tracks the echo path between the loudspeaker output and the microphone. During training, the network and filter are jointly optimized to learn the network parameters. This system requires 17 thousand parameters that consume 500 Million floating-point operations per second and 40 Kilo-bytes of memory. It also satisfies hands-free communication timing requirements on a standard neural processor, which renders it adequate for embedding on hands-free communication devices. Using 280 hours of real and synthetic data, experiments show advantageous performance compared to competing methods.
Amir Ivry, Israel Cohen, Baruch Berdugo
Interspeech3
2021 Controlling Elevation and Azimuth Beamwidths With Concentric Circular Microphone Arrays
abstract
Solutions for frequency-invariant or constant-beamwidth beamforming with concentric circular arrays (CCAs) generally control only the azimuth-beampattern. This work introduces a beamforming methodology, which instead of designing a fixed beamwidth across the spectrum, maximizes the directivity-factor under the constraint that the beamwidth is greater than or equal to a pre-defined target-beamwidth, in both the elevation and the azimuth. The principal idea lies in dynamically weighting the microphones of the CCA based on two factors - (i) their corresponding ring number or radius, and (ii) their alignment with the signal-path. The parameters controlling the weights are determined using a modified gradient-descent algorithm. Experimental results show the flexibility and advantage of the proposed methodology over state-of-the-art methods.
Rajib Sharma, Israel Cohen, Baruch Berdugo
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Evaluation of Deep-Learning-Based Voice Activity Detectors and Room Impulse Response Models in Reverberant Environments
abstract
State-of-the-art deep-learning-based voice activity detectors (VADs) are often trained with anechoic data. However, real acoustic environments are generally reverberant, which causes the performance to significantly deteriorate. To mitigate this mismatch between training data and real data, we simulate an augmented training set that contains nearly five million utterances. This extension comprises of anechoic utterances and their reverberant modifications, generated by convolutions of the anechoic utterances with a variety of room impulse responses (RIRs). We consider five different models to generate RIRs, and five different VADs that are trained with the augmented training set. We test all trained systems in three different real reverberant environments. Experimental results show 20% increase on average in accuracy, precision and recall for all detectors and response models, compared to anechoic training. Furthermore, one of the RIR models consistently yields better performance than the other models, for all the tested VADs. Additionally, one of the VADs consistently outperformed the other VADs in all experiments.
Amir Ivry, Israel Cohen, Baruch Berdugo
ICASSP3
2003 Two-channel signal detection and speech enhancement based on the transient beam-to-reference ratio
abstract
In reverberant and noisy environments, multichannel systems are designed for spatially filtering interfering signals coming from undesired directions. In case of incoherent or diffuse noise fields, beamforming alone does not provide sufficient noise reduction, and post-filtering is normally required. In this paper, we present a two-channel post-filtering approach for signal detection and speech enhancement. A mild assumption is made, that a desired signal component is stronger at the beamformer output than at the reference noise signal, and a noise component is stronger at the reference signal. The ratio between the transient power at the beamformer output and the transient power at the reference noise signal is used for indicating whether such a transient is desired or interfering. Experimental results demonstrate the usefulness of the proposed approach in a car environment.
Israel Cohen, Baruch Berdugo
ICASSP (5)2
2003 Multichannel signal detection based on the transient beam-to-reference ratio
abstract
We present a multichannel signal detection approach that is particularly advantageous in nonstationary noise environments. A beamformer is realistically assumed to have a steering error, a blocking matrix that is unable to block all of the desired signal components, and a noise canceler that is adapted to the pseudostationary noise, but not modified during transient interference. Signal components are detected at the beamformer output based on a measure of their local nonstationarity, and discriminated from transient noise components based on the transient beam-to-reference ratio.
Israel Cohen, Baruch Berdugo
IEEE Signal Process. Lett.2
2002 Microphone array post-filtering for non-stationary noise suppression
abstract
Microphone array post-filtering allows additional reduction of noise components at a beamformer output. Existing techniques are either restricted to classical delay-and-sum beamformers, or are based on single-channel speech enhancement algorithms that are inefficient at attenuating highly non-stationary noise components. In this paper, we introduce a microphone array post-filtering approach, applicable to adaptive beamformer, that differentiates non-stationary noise components from speech components. The ratio between the transient power at the beamformer primary output and the transient power at the reference noise signals is used for indicating whether such a transient is desired or interfering. Based on a Gaussian statistical model and combined with an appropriate spectral enhancement technique, a significantly reduced level of non-stationary noise is achieved without further distorting speech components. Experimental results demonstrate the effectiveness of the proposed method.
Israel Cohen, Baruch Berdugo
ICASSP2
2002 Speakers' direction finding using estimated time delays in the frequency domain
Baruch Berdugo, Judith Rosenhouse, Haim Azhari
Signal Process.1
2002 Noise estimation by minima controlled recursive averaging for robust speech enhancement
abstract
In this letter, we introduce a minima controlled recursive averaging (MCRA) approach for noise estimation. The noise estimate is given by averaging past spectral power values and using a smoothing parameter that is adjusted by the signal presence probability in subbands. The presence of speech in subbands is determined by the ratio between the local energy of the noisy speech and its minimum within a specified time window. The noise estimate is computationally efficient, robust with respect to the input signal-to-noise ratio (SNR) and type of underlying additive noise, and characterized by the ability to quickly follow abrupt changes in the noise spectrum.
Israel Cohen, Baruch Berdugo
IEEE Signal Process. Lett.2
2001 Speech enhancement for non-stationary noise environments
Israel Cohen, Baruch Berdugo
Signal Process.2