EDBT 2026 Demo / reviewers in the wild / expert
Eliya Nachmani
dblp:183/6370
· DBLP profile ↗
22ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0003-4220-5672ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SimulTron: On-Device Simultaneous Speech to Speech TranslationabstractSimultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST architecture designed to tackle this task. SimulTron is a lightweight direct S2ST model that uses the strengths of the Translatotron framework while incorporating key modifications for streaming operation, and an adjustable fixed delay. Our experiments show that SimulTron surpasses Translatotron 2 in offline evaluations. Furthermore, real-time evaluations reveal that SimulTron improves upon the performance achieved by Translatotron 1. Additionally, SimulTron achieves superior BLEU scores and latency compared to previous real-time S2ST method on the MuST-C dataset. Significantly, we have successfully deployed SimulTron on a Pixel 7 Pro device, show its potential for simultaneous S2ST on-device. Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding 0004, Ye Jia, Nadav Bar, Heiga Zen, Michelle Tadmor Ramanovich |
ICASSP | 2 |
| 2025 | Zero-Shot Mono-to-Binaural Speech Synthesis
Alon Levkovitch, Julian Salazar, Soroosh Mariooryad, R. J. Skerry-Ryan, Nadav Bar, W. Bastiaan Kleijn, Eliya Nachmani |
INTERSPEECH | 7 |
| 2024 | Translatotron 3: Speech to Speech Translation with Monolingual DataabstractThis paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation. Experimental results in speech-to-speech translation tasks between Spanish and English show that Translatotron 3 outperforms a baseline cascade system, reporting 18.14 BLEU points improvement on the synthesized Unpaired-Conversational dataset. In contrast to supervised approaches that necessitate real paired data, or specialized modeling to replicate para-/non-linguistic information such as pauses, speaking rates, and speaker identity, Translatotron 3 showcases its capability to retain it. Eliya Nachmani, Alon Levkovitch, Yifan Ding 0004, Chulayuth Asawaroengchai, Heiga Zen, Michelle Tadmor Ramanovich |
ICASSP | 1 |
| 2024 | Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source SeparationabstractThe problem of speech separation, also known as the cocktail party problem,
refers to the task of isolating a single speech signal from a mixture of speech
signals. Previous work on source separation derived an upper bound for the
source separation task in the domain of human speech. This bound is derived for
deterministic models. Recent advancements in generative models challenge this
bound. We show how the upper bound can be generalized to the case of random
generative models. Applying a diffusion model Vocoder that was pretrained to
model single-speaker voices on the output of a deterministic separation model leads
to state-of-the-art separation results. It is shown that this requires one to combine
the output of the separation model with that of the diffusion model. In our method,
a linear combination is performed, in the frequency domain, using weights that are
inferred by a learned model. We show state-of-the-art results on 2, 3, 5, 10, and 20
speakers on multiple benchmarks. In particular, for two speakers, our method is
able to surpass what was previously considered the upper performance bound. Shahar Lutati, Eliya Nachmani, Lior Wolf |
ICLR | 2 |
| 2024 | Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMabstractWe present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire system is trained end-to-end and operates directly on spectrograms, simplifying our architecture. Key to our approach is a training objective that jointly supervises speech recognition, text continuation, and speech synthesis using only paired speech-text pairs, enabling a `cross-modal' chain-of-thought within a single decoding pass. Our method surpasses existing spoken language models in speaker preservation and semantic coherence. Furthermore, the proposed model improves upon direct initialization in retaining the knowledge of the original LLM as demonstrated through spoken QA datasets. We release our audio samples and spoken QA dataset via our website. Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaroengchai, Soroosh Mariooryad, Ehud Rivlin, R. J. Skerry-Ryan, Michelle Tadmor Ramanovich |
ICLR | 1 |
| 2024 | Harnessing the flexibility of neural networks to predict dynamic theoretical parameters underlying human choice behaviorabstractReinforcement learning (RL) models are used extensively to study human behavior. These rely on normative models of behavior and stress interpretability over predictive capabilities. More recently, neural network models have emerged as a descriptive modeling paradigm that is capable of high predictive power yet with limited interpretability. Here, we seek to augment the expressiveness of theoretical RL models with the high flexibility and predictive power of neural networks. We introduce a novel framework, which we term theoretical-RNN (t-RNN), whereby a recurrent neural network is trained to predict trial-by-trial behavior and to infer theoretical RL parameters using artificial data of RL agents performing a two-armed bandit task. In three studies, we then examined the use of our approach to dynamically predict unseen behavior along with time-varying theoretical RL parameters. We first validate our approach using synthetic data with known RL parameters. Next, as a proof-of-concept, we applied our framework to two independent datasets of humans performing the same task. In the first dataset, we describe differences in theoretical RL parameters dynamic among clinical psychiatric vs. healthy controls. In the second dataset, we show that the exploration strategies of humans varied dynamically in response to task phase and difficulty. For all analyses, we found better performance in the prediction of actions for t-RNN compared to the stationary maximum-likelihood RL method. We discuss the use of neural networks to facilitate the estimation of latent RL parameters underlying choice behavior. Yoav Ger, Eliya Nachmani, Lior Wolf, Nitzan Shahar |
PLoS Comput. Biol. | 2 |
| 2023 | Decision S4: Efficient Sequence-Based RL via State Spaces Layers
Shmuel Bar-David, Itamar Zimerman, Eliya Nachmani, Lior Wolf |
ICLR | 3 |
| 2023 | kNN-Diffusion: Image Generation via Large-Scale Retrieval
Shelly Sheynin, Oron Ashual, Adam Polyak, Uriel Singer, Oran Gafni, Eliya Nachmani, Yaniv Taigman |
ICLR | 6 |
| 2022 | Zero-Shot Voice Conditioning for Denoising Diffusion TTS ModelsabstractWe present a novel way of conditioning a pretrained denoising diffusion speech model to produce speech in the voice of a novel person unseen during training.The method requires a short (∼ 3 seconds) sample from the target person, and generation is steered at inference time, without any training steps.At the heart of the method lies a sampling process that combines the estimation of the denoising model with a low-pass version of the new speaker's sample.The objective and subjective evaluations show that our sampling method can generate a voice similar to that of the target speaker in terms of frequency, with an accuracy comparable to state-of-the-art methods, and without training. Alon Levkovitch, Eliya Nachmani, Lior Wolf |
INTERSPEECH | 2 |
| 2022 | SepIt: Approaching a Single Channel Speech Separation BoundabstractWe present an upper bound for the Single Channel Speech Separation task, which is based on an assumption regarding the nature of short segments of speech.Using the bound, we are able to show that while the recent methods have made great progress for a few speakers, there is room for improvement for five and ten speakers.We then introduce a Deep neural network, SepIt, that iteratively improves the different speakers' estimation.At test time, SpeIt has a varying number of iterations per test sample, based on a mutual information criterion that arises from our analysis.In an extensive set of experiments, SepIt outperforms the state of the art neural networks for 2, 3, 5, and 10 speakers. Shahar Lutati, Eliya Nachmani, Lior Wolf |
INTERSPEECH | 2 |
| 2022 | A-Muze-Net: Music Generation by Composing the Harmony Based on the Generated Melody
Or Goren, Eliya Nachmani, Lior Wolf |
MMM (1) | 2 |
| 2021 | Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy SettingsabstractWe present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the time domain, together with parameter sharing between all separation heads. The classification branch estimates the number of speakers while each head is specialized in separating a different number of speakers. We evaluate the proposed model under both clean and noisy reverberant settings. Results suggest that the proposed approach is superior to the baseline model by a significant margin. Additionally, we present a new noisy and reverberant dataset of up to five different speakers speaking simultaneously. Shlomo E. Chazan, Lior Wolf, Eliya Nachmani, Yossi Adi |
ICASSP | 3 |
| 2021 | Recovering AES Keys with a Deep Cold Boot AttackabstractCold boot attacks inspect the corrupted random access memory soon after the power has been shut down. While most of the bits have been corrupted, many bits, at random locations, have not. Since the keys in many encryption schemes are being expanded in memory into longer keys with fixed redundancies, the keys can often be restored. In this work we combine a deep error correcting code technique together with a modified SAT solver scheme in order to apply the attack to AES keys. Even though AES consists Rijndael SBOX elements, that are specifically designed to be resistant to linear and differential cryptanalysis, our method provides a novel formalization of the AES key scheduling as a computational graph, which is implemented by neural message passing network. Our results show that our methods outperform the state of the art attack methods by a very large gap. Itamar Zimerman, Eliya Nachmani, Lior Wolf |
ICML | 2 |
| 2021 | Many-Speakers Single Channel Speech Separation with Optimal Permutation TrainingabstractSingle channel speech separation has experienced great progress in the last few years. However, training neural speech separation for a large number of speakers (e.g., more than 10 speakers) is out of reach for the current methods, which rely on the Permutation Invariant Loss (PIT). In this work, we present a permutation invariant training that employs the Hungarian algorithm in order to train with an $O(C^3)$ time complexity, where $C$ is the number of speakers, in comparison to $O(C!)$ of PIT based methods. Furthermore, we present a modified architecture that can handle the increased number of speakers. Our approach separates up to $20$ speakers and improves the previous results for large $C$ by a wide margin. Shaked Dovrat, Eliya Nachmani, Lior Wolf |
Interspeech | 2 |
| 2021 | SAGRNN: Self-Attentive Gated RNN For Binaural Speaker Separation With Interaural Cue PreservationabstractMost existing deep learning based binaural speaker separation systems focus on producing a monaural estimate for each of the target speakers, and thus do not preserve the interaural cues, which are crucial for human listeners to perform sound localization and lateralization. In this study, we address talker-independent binaural speaker separation with interaural cues preserved in the estimated binaural signals. Specifically, we extend a newly-developed gated recurrent neural network for monaural separation by additionally incorporating self-attention mechanisms and dense connectivity. We develop an end-to-end multiple-input multiple-output system, which directly maps from the binaural waveform of the mixture to those of the speech signals. The experimental results show that our proposed approach achieves significantly better separation performance than a recent binaural separation approach. In addition, our approach effectively preserves the interaural cues, which improves the accuracy of sound localization. Ke Tan 0001, Buye Xu, Anurag Kumar 0003, Eliya Nachmani, Yossi Adi |
IEEE Signal Process. Lett. | 4 |
| 2020 | A Gated Hypernet Decoder for Polar CodesabstractHypernetworks were recently shown to improve the performance of message passing algorithms for decoding error correcting codes. In this work, we demonstrate how hypernet-works can be applied to decode polar codes by employing a new formalization of the polar belief propagation decoding scheme. We demonstrate that our method improves the previous results of neural polar decoders and achieves, for large SNRs, the same bit-error-rate performances as the successive list cancellation method, which is known to be better than any belief propagation decoders and very close to the maximum likelihood decoder. Eliya Nachmani, Lior Wolf |
ICASSP | 1 |
| 2020 | Voice Separation with an Unknown Number of Multiple SpeakersabstractWe present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously. The new method employs gated neural networks that are trained to separate the voices at multiple processing steps, while maintaining the speaker in each output channel fixed. A different model is trained for every number of possible speakers, and the model with the largest number of speakers is employed to select the actual number of speakers in a given sample. Our method greatly outperforms the current state of the art, which, as we show, is not competitive for more than two speakers. Eliya Nachmani, Yossi Adi, Lior Wolf |
ICML | 1 |
| 2019 | Unsupervised Polyglot Text-to-speechabstractWe present a TTS neural network that is able to produce speech in multiple languages. The proposed network is able to transfer a voice, which was presented as a sample in a source language, into one of several target languages. Training is done without using matching or parallel data, i.e., without samples of the same speaker in multiple languages, making the method much more applicable. The conversion is based on learning a polyglot network that has multiple per-language sub-networks and adding loss terms that preserve the speaker's identity in multiple languages. We evaluate the proposed polyglot neural network for three languages with a total of more than 400 speakers and demonstrate convincing conversion capabilities. Eliya Nachmani, Lior Wolf |
ICASSP | 1 |
| 2019 | Unsupervised Singing Voice ConversionabstractWe present a deep learning method for singing voice conversion. The proposed network is not conditioned on the text or on the notes, and it directly converts the audio of one singer to the voice of another. Training is performed without any form of supervision: no lyrics or any kind of phonetic features, no notes, and no matching samples between singers. The proposed network employs a single CNN encoder for all singers, a single WaveNet decoder, and a classifier that enforces the latent representation to be singer-agnostic. Each singer is represented by one embedding vector, which the decoder is conditioned on. In order to deal with relatively small datasets, we propose a new data augmentation scheme, as well as new training losses and protocols that are based on backtranslation. Our evaluation presents evidence that the conversion produces natural signing voices that are highly recognizable as the target singer. Eliya Nachmani, Lior Wolf |
INTERSPEECH | 1 |
| 2019 | Hyper-Graph-Network Decoders for Block CodesabstractNeural decoders were shown to outperform classical message passing techniques for short BCH codes. In this work, we extend these results to much larger families of algebraic block codes, by performing message passing with graph neural networks. The parameters of the sub-network at each variable-node in the Tanner graph are obtained from a hypernetwork that receives the absolute values of the current message as input. To add stability, we employ a simplified version of the arctanh activation that is based on a high order Taylor approximation of this activation function. Our results show that for a large number of algebraic block codes, from diverse families of codes (BCH, LDPC, Polar), the decoding obtained with our method outperforms the vanilla belief propagation method as well as other learning techniques from the literature. Eliya Nachmani, Lior Wolf |
NeurIPS | 1 |
| 2018 | VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop
Yaniv Taigman, Lior Wolf, Adam Polyak, Eliya Nachmani |
ICLR (Poster) | 4 |
| 2018 | Fitting New Speakers Based on a Short Untranscribed SampleabstractLearning-based Text To Speech systems have the potential to generalize from one speaker to the next and thus require a relatively short sample of any new voice. However, this promise is currently largely unrealized. We present a method that is designed to capture a new speaker from a short untranscribed audio sample. This is done by employing an additional network that given an audio sample, places the speaker in the embedding space. This network is trained as part of the speech synthesis system using various consistency losses. Our results demonstrate a greatly improved performance on both the dataset speakers, and, more importantly, when fitting new voices, even from very short samples. Eliya Nachmani, Adam Polyak, Yaniv Taigman, Lior Wolf |
ICML | 1 |