VLDB 2026 Research / reviewers in the wild / expert
Amir Ivry
dblp:241/3569
· DBLP profile ↗
10ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0002-9043-7040ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | E-URES 2.0: Efficient User-Centric Residual-Echo Suppression with a Lightweight Neural NetworkabstractWe recently introduced the Efficient User-centric Residual-Echo Suppression (E-URES) framework, which significantly reduces the floating-point operations per second (FLOPS) required during inference by 90% compared to the URES framework. The E-URES operates based on a user-operating point (UOP) defined by two key metrics: the residual echo suppression level (RESL) and the desired-speech maintained level (DSML) that the user anticipates from the output signal of a residual echo suppression (RES) system. In the first stage, an ensemble of 101 branches is employed, where each branch has two cascaded neural networks: a preliminary RES system with a design parameter, which varies between branches and balances the RESL and DSML of its RES systems’ prediction, and a subsequent UOP estimator. In the second stage, a neural network uses available acoustic signals and the UOP to predict which three branches achieve the highest acoustic echo cancellation mean opinion score (AECMOS) within a specified UOP-error tolerance. Then, costly AECMOS calculations are performed only for these selected branches. Despite this efficiency mechanism, the E-URES can apply real-time inference only with dedicated and expensive hardware, limiting its wide adoption. Here, we present E-URES 2.0, which focuses on reducing the computational costs of E-URES in its first stage. A lightweight neural network preprocesses available acoustic signals and the UOP to track a subset of the 101 design parameters that their branches produce the most accurate UOP estimations in their outcomes. Only these branches are calculated during inference and continue to the AECMOS estimation stage. With 60 hours of data, we show that with a negligible performance drop on average, the E-URES 2.0 can reduce 87% of the branches and 61% of the FLOPS of the E-URES and can achieve real-time inference with standard, affordable hardware. Amir Ivry, Israel Cohen |
ICASSP | 1 |
| 2025 | Summary of the NOTSOFAR-1 challenge: Highlights and learnings
Igor Abramovski, Alon Vinnikov, Shalev Shaer, Naoyuki Kanda, Xiaofei Wang 0009, Amir Ivry, Eyal Krupka |
Comput. Speech Lang. | 6 |
| 2024 | NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
Alon Vinnikov, Amir Ivry, Aviv Hurvitz, Igor Abramovski, Sharon Koubi, Ilya Gurvich, Shai Pe'er, Benjamin Elizalde, Naoyuki Kanda, Xiaofei Wang 0009, Shalev Shaer, Stav Yagev, Yossi Asher, Sunit Sivasankaran, Yifan Gong 0001, Huaming Wang, Eyal Krupka |
INTERSPEECH | 2 |
| 2024 | A User-Centric Approach for Deep Residual-Echo Suppression in Double-TalkabstractWe introduce a user-centric residual-echo suppression (URES) framework in double-talk. This framework receives a user operating point (UOP) that consists of two metric values: the residual echo suppression level (RESL) and the desired speech-maintained level (DSML) that the user expects from the RES outcome. Then, the URES pipeline undergoes three stages. Firstly, we consider a deep RES model with a tunable design parameter that balances between the RESL and DSML and utilizes 101 pre-trained instances of this model, each with a different design parameter value. Thus, an identical input is expected to generate a different pair of RESL and DSML values in the prediction of every instance. Second, every prediction is separately fed to a subsequent pre-trained deep model instance that estimates the RESL and DSML of the prediction since these metrics depend on unavailable information in practice. Lastly, each pair of RESL and DSML estimates is compared with the UOP. The pairs that match the UOP up to a given tolerance threshold are narrowed down to the prediction with the maximal acoustic-echo cancellation mean-opinion score (AECMOS), which is the output of the URES system. This suggested framework holds three prominent advantages introduced in this study: it generates an RES output with RESL and DSML that match a UOP, supports near-real-time tracking of UOP changes, and applies AECMOS maximization. Experimental results consider 60 h of varied real and synthetic data. Average results can achieve an AECMOS subjectively considered excellent with RESL and DSML deviations of roughly 2 dB from the UOP. Any UOP adjustment can be tracked in less than 40 ms with a real-time factor of 1.92, but due to the high computational resources demanded by the framework, this is enabled on-edge only with high-end dedicated hardware, which limits general availability. Amir Ivry, Israel Cohen, Baruch Berdugo |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Deep Adaptation Control for Acoustic Echo CancellationabstractWe propose a general framework for adaptation control using deep neural networks (NNs) and apply it to acoustic echo cancellation (AEC). First, the optimal step-size that controls the adaptation is derived offline by solving a constrained nonlinear optimization problem that minimizes the adaptive filter misadjustment. Then, a deep NN is trained to learn the relation between the input data and the optimal step-size. In real-time, the NN infers the optimal step-size from streaming data and feeds it to an NLMS filter for AEC. This data-driven method makes no assumptions on the acoustic setup and is entirely non-parametric. Experiments with 100 h of real and synthetic data show that the proposed method outperforms the competition in echo cancellation, speech distortion, and convergence during both single-talk and double-talk. Amir Ivry, Israel Cohen, Baruch Berdugo |
ICASSP | 1 |
| 2022 | Off-the-Shelf Deep Integration For Residual-Echo SuppressionabstractResidual-echo suppression (RES) systems suppress the echo and preserve the speech from a mixture of the two. In hands-free speech communication, RES may also be addressed as a source separation (SS) or speech enhancement (SE) problem, where the echo can be manipulated as an interfering speech signal. In this study, we fine-tune three pre-trained deep learning-based systems originally designed for RES, SS, and SE, and show that the best performing system for the task of RES varies with respect to the acoustic conditions. Then, we propose a real-time data-driven integration of these systems, where a neural network continuously tracks the system that achieves the best performance during both single-talk and double-talk periods. Experiments with 100 h of real and synthetic data show that the integrated system outperforms each individual system in terms of echo suppression and speech distortion in various acoustic environments. Amir Ivry, Israel Cohen, Baruch Berdugo |
ICASSP | 1 |
| 2022 | Objective Metrics to Evaluate Residual-Echo Suppression During Double-Talk in the Stereophonic Case
Amir Ivry, Israel Cohen, Baruch Berdugo |
INTERSPEECH | 1 |
| 2021 | Deep Residual Echo Suppression With A Tunable Tradeoff Between Signal Distortion And Echo SuppressionabstractIn this paper, we propose a residual echo suppression method using a UNet neural network that directly maps the outputs of a linear acoustic echo canceler to the desired signal in the spectral domain. This system embeds a design parameter that allows a tunable tradeoff between the desired-signal distortion and residual echo suppression in double-talk scenarios. The system employs 136 thousand parameters, and requires 1.6 Giga floating-point operations per second and 10 Mega-bytes of memory. The implementation satisfies both the timing requirements of the AEC challenge and the computational and memory limitations of on-device applications. Experiments are conducted with 161 h of data from the AEC challenge database and from real independent recordings. We demonstrate the performance of the proposed system in real-life conditions and compare it with two competing methods regarding echo suppression and desired-signal distortion, generalization to various environments, and robustness to high echo levels. Amir Ivry, Israel Cohen, Baruch Berdugo |
ICASSP | 1 |
| 2021 | Nonlinear Acoustic Echo Cancellation with Deep LearningabstractWe propose a nonlinear acoustic echo cancellation system, which aims to model the echo path from the far-end signal to the near-end microphone in two parts. Inspired by the physical behavior of modern hands-free devices, we first introduce a novel neural network architecture that is specifically designed to model the nonlinear distortions these devices induce between receiving and playing the far-end signal. To account for variations between devices, we construct this network with trainable memory length and nonlinear activation functions that are not parameterized in advance, but are rather optimized during the training stage using the training data. Second, the network is succeeded by a standard adaptive linear filter that constantly tracks the echo path between the loudspeaker output and the microphone. During training, the network and filter are jointly optimized to learn the network parameters. This system requires 17 thousand parameters that consume 500 Million floating-point operations per second and 40 Kilo-bytes of memory. It also satisfies hands-free communication timing requirements on a standard neural processor, which renders it adequate for embedding on hands-free communication devices. Using 280 hours of real and synthetic data, experiments show advantageous performance compared to competing methods. Amir Ivry, Israel Cohen, Baruch Berdugo |
Interspeech | 1 |
| 2020 | Evaluation of Deep-Learning-Based Voice Activity Detectors and Room Impulse Response Models in Reverberant EnvironmentsabstractState-of-the-art deep-learning-based voice activity detectors (VADs) are often trained with anechoic data. However, real acoustic environments are generally reverberant, which causes the performance to significantly deteriorate. To mitigate this mismatch between training data and real data, we simulate an augmented training set that contains nearly five million utterances. This extension comprises of anechoic utterances and their reverberant modifications, generated by convolutions of the anechoic utterances with a variety of room impulse responses (RIRs). We consider five different models to generate RIRs, and five different VADs that are trained with the augmented training set. We test all trained systems in three different real reverberant environments. Experimental results show 20% increase on average in accuracy, precision and recall for all detectors and response models, compared to anechoic training. Furthermore, one of the RIR models consistently yields better performance than the other models, for all the tested VADs. Additionally, one of the VADs consistently outperformed the other VADs in all experiments. Amir Ivry, Israel Cohen, Baruch Berdugo |
ICASSP | 1 |