Israel D. Gebru

dblp:154/4306 · also Israel Dejene Gebru · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
abstract
We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods.
Mingfei Chen, Israel D. Gebru, Ishwarya Ananthabhotla, Christian Richardt, Dejan Markovic, Jake Sandakly, Steven Krenn, Todd Keebler, Eli Shlizerman, Alexander Richard
CVPR2
2025 A2B: Neural Rendering of Ambisonic Recordings to Binaural
abstract
This paper introduces a novel neural network model for rendering binaural audio directly from ambisonic recordings. We optimized the model end-to-end to learn a direct mapping between ambisonic and binaural signals. Our approach eliminates traditional processing steps that were required to mitigate artifacts due to spherical harmonic order truncation and spatial aliasing, as well as other complex filtering needed to compensate for near-field sound sources. To showcase the advantage of neural network-based rendering over traditional signal processing approaches, we introduce a new dataset that includes challenging near-field sound sources, including speech and background noises. We demonstrate that our model can produce binaural audio results that closely match the fidelity of ground truth binaural recordings. Our comprehensive validation shows that the proposed method outperforms existing methods on several error metrics as well as in subjective evaluations. Model code, demos and datasets are available on our project webpage.
Israel D. Gebru, Todd Keebler, Jake Sandakly, Steven Krenn, Dejan Markovic, Julia Buffalini, Samuel Hassel, Alexander Richard
ICASSP1
2025 ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling
abstract
Neural audio codecs have been widely adopted in audio-generative tasks because their compact and discrete representations are suitable for both large-language-model-style and regression-based generative models. However, most neural codecs struggle to model out-of-domain audio, resulting in error propagations to downstream generative tasks. In this paper, we first argue that information loss from codec compression degrades out-of-domain robustness. Then, we propose full-band 48 kHz ComplexDec with complex spectral input and output to ease the information loss while adopting the same 24 kbps bitrate as the baseline AuidoDec and ScoreDec. Objective and subjective evaluations demonstrate the out-of-domain robustness of ComplexDec trained using only the 30-hour VCTK corpus.
Yi-Chiao Wu, Dejan Markovic, Steven Krenn, Israel D. Gebru, Alexander Richard
ICASSP4
2025 BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models
abstract
Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-quality binaural audio that is indistinguishable from real-world recordings requires precise modeling of binaural cues, room reverb, and ambient sounds. Additionally, real-world applications demand streaming inference. To address these challenges, we propose a flow matching based streaming binaural speech synthesis framework called BinauralFlow. We consider binaural rendering to be a generation problem rather than a regression problem and design a conditional flow matching model to render high-quality audio. Moreover, we design a causal U-Net architecture that estimates the current audio frame solely based on past information to tailor generative models for streaming inference. Finally, we introduce a continuous inference pipeline incorporating streaming STFT/ISTFT operations, a buffer bank, a midpoint solver, and an early skip schedule to improve rendering continuity and speed. Quantitative and qualitative evaluations demonstrate the superiority of our method over SOTA approaches. A perceptual study further reveals that our model is nearly indistinguishable from real-world recordings, with a 42% confusion rate.
Susan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn, Todd Keebler, Jake Sandakly, Frank Yu, Samuel Hassel, Chenliang Xu, Alexander Richard
ICML3
2024 Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
abstract
We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modali-ties. The dataset includes high-quality and densely captured room impulse response data paired with multi-view images, and precise 6DoF pose tracking data for sound emitters and listeners in the rooms. We used this dataset to evaluate existing methods for novel-view acoustic synthesis and impulse re-sponse generation which previously relied on synthetic data. In our evaluation, we thoroughly assessed existing audio and audio- visual models against multiple criteria and proposed settings to enhance their performance on real-world data. We also conducted experiments to investigate the impact of incorporating visual data (i.e., images and depth) into neu-ral acoustic field models. Additionally, we demonstrated the effectiveness of a simple sim2real approach, where a model is pre-trained with simulated data and fine-tuned with sparse real-world data, resulting in significant improvements in the few-shot learning approach. RAF is the first dataset to provide densely captured room acoustic data, making it an ideal resource for researchers working on audio and audio-visual neural acoustic field modeling techniques. Demos and datasets are available on our project page.
Israel D. Gebru, Christian Richardt, Anurag Kumar 0003, William Laney, Andrew Owens, Alexander Richard
CVPR2
2024 ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-Filter
abstract
Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative adversarial network (GAN) training is usually required for these E2E neural codecs because of the difficulty of direct phase modeling. However, such adversarial learning hinders these codecs from preserving the original phase information. To achieve human-level naturalness with a reasonable bitrate, preserve the original phase, and get rid of the tricky and opaque GAN training, we develop a score-based diffusion post-filter (SPF) in the complex spectral domain and combine our previous AudioDec with the SPF to propose ScoreDec, which can be trained using only spectral and score-matching losses. Both the objective and subjective experimental results show that ScoreDec with a 24 kbps bitrate encodes and decodes full-band 48 kHz speech with human-level naturalness and well-preserved phase information.
Yi-Chiao Wu, Dejan Markovic, Steven Krenn, Israel D. Gebru, Alexander Richard
ICASSP4
2023 Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural Speech
abstract
We propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer to the listener’s head), and quantification tasks (i.e. quantifying the depth difference between the inputs). In addition, training leverages recent advances in metric and multi-task learning, which allows the framework to be invariant to both signal content (i.e. non-matched reference) and directional cues (i.e. azimuth and elevation). Our framework has additional useful qualities that make it suitable for use as an objective metric to benchmark binaural audio systems, particularly depth perception and sound externalization, which we demonstrate through experiments. We also show that NORD generalizes well under different reverberation and environments. The results from preference and quantification tasks correlate well with measured results.
Pranay Manocha, Israel D. Gebru, Anurag Kumar 0003, Dejan Markovic, Alexander Richard
ICASSP2
2023 Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec
abstract
A good audio codec for live applications such as telecommunication is characterized by three key properties: (1) compression, i.e. the bitrate that is required to transmit the signal should be as low as possible; (2) latency, i.e. encoding and decoding the signal needs to be fast enough to enable communication without or with only minimal noticeable delay; and (3) reconstruction quality of the signal. In this work, we propose an open-source, streamable, and real-time neural audio codec that achieves strong performance along all three axes: it can reconstruct highly natural sounding 48 kHz speech signals while operating at only 12 kbps and running with less than 6 ms (GPU)/10 ms (CPU) latency. An efficient training paradigm is also demonstrated for developing such neural audio codecs for real-world scenarios. Both objective and subjective evaluations using the VCTK corpus are provided. To sum up, AudioDec is a well-developed plug-and-play benchmark for audio codec applications.
Yi-Chiao Wu, Israel D. Gebru, Dejan Markovic, Alexander Richard
ICASSP2
2023 Spatialization Quality Metric for Binaural Speech
Pranay Manocha, Israel D. Gebru, Anurag Kumar 0003, Dejan Markovic, Alexander Richard
INTERSPEECH2
2022 End-to-End Binaural Speech Synthesis
abstract
In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing environmental factors like ambient noise or reverb.The network is a modified vectorquantized variational autoencoder, trained with several carefully designed objectives, including an adversarial loss.We evaluate the proposed system on an internal binaural dataset with objective metrics and a perceptual study.Results show that the proposed approach matches the ground truth data more closely than previous methods.In particular, we demonstrate the capability of the adversarial loss in capturing environment effects needed to create an authentic auditory scene.
Wen-Chin Huang, Dejan Markovic, Alexander Richard, Israel D. Gebru, Anjali Menon
INTERSPEECH4
2022 SAQAM: Spatial Audio Quality Assessment Metric
abstract
Audio quality assessment is critical for assessing the perceptual realism of sounds.However, the time and expense of obtaining "gold standard" human judgments limit the availability of such data.For AR&VR, good perceived sound quality and localizability of sources are among the key elements to ensure complete immersion of the user.Our work introduces SAQAM which uses a multi-task learning framework to assess listening quality (LQ) and spatialization quality (SQ) between any given pair of binaural signals without using any subjective data.We model LQ by training on a simulated dataset of triplet human judgments, and SQ by utilizing activation-level distances from networks trained for direction of arrival (DOA) estimation.We show that SAQAM correlates well with human responses across four diverse datasets.Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks.For example, simply replacing an existing loss with our metric yields improvement in a speech-enhancement network.
Pranay Manocha, Anurag Kumar 0003, Buye Xu, Anjali Menon, Israel D. Gebru, Vamsi K. Ithapu, Paul Calamia
INTERSPEECH5
2021 Implicit HRTF Modeling Using Temporal Convolutional Networks
abstract
Estimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at thousands of distinct spatial directions in an anechoic chamber. In this work, we present a data-driven approach to learn HRTFs implicitly with a neural network that achieves state of the art results compared to traditional approaches but relies on a much simpler data capture that can be performed in arbitrary, non-anechoic rooms. Despite that simpler and less acoustically ideal data capture, our deep learning based approach learns HRTF of high quality. We show in a perceptual study that the produced binaural audio is ranked on par with traditional DSP approaches by humans and illustrate that interaural time differences (ITDs), interaural level differences (ILDs) and spectral clues are accurately estimated.
Israel D. Gebru, Dejan Markovic, Alexander Richard, Steven Krenn, Gladstone Alexander Butler, Fernando De la Torre, Yaser Sheikh
ICASSP1
2021 Neural Synthesis of Binaural Speech From Mono Audio
Alexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn, Gladstone Alexander Butler, Fernando De la Torre, Yaser Sheikh
ICLR3
2019 Soundfield Reconstruction in Reverberant Environments Using Higher-order Microphones and Impulse Response Measurements
abstract
This paper addresses the problem of soundfield reconstruction over a large area using a distributed array of higher-order microphones. Given an area enclosed by the array, one can distinguish between two components of the soundfield: the interior soundfield generated by sources outside of the enclosed area and the exterior soundfield generated by sources inside the enclosed area. These components form an indistinguishable mixture and, despite the existence of theoretical solutions to separate them, practical implementation is challenging due to high number of microphones needed for large regions and high frequencies. In this work, we consider a scenario where the interior soundfield is characterized by reverberation and show how a set of RIR measurements can be used to parametrize the interior component as a function of the exterior component, effectively reducing the unknowns of the problem.
Federico Borra, Israel D. Gebru, Dejan Markovic
ICASSP2
2018 Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion
abstract
Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants engaged in multi-party interaction while they move around and turn their heads towards the other participants rather than facing the cameras and the microphones. Multiple-person visual tracking is combined with multiple speech-source localization in order to tackle the speech-to-person association problem. The latter is solved within a novel audio-visual fusion method on the following grounds: binaural spectral features are first extracted from a microphone pair, then a supervised audio-visual alignment technique maps these features onto an image, and finally a semi-supervised clustering method assigns binaural spectral features to visible persons. The main advantage of this method over previous work is that it processes in a principled way speech signals uttered simultaneously by multiple persons. The diarization itself is cast into a latent-variable temporal graphical model that infers speaker identities and speech turns, based on the output of an audio-visual association process, executed at each time slice, and on the dynamics of the diarization variable itself. The proposed formulation yields an efficient exact inference procedure. A novel dataset, that contains audio-visual training data as well as a number of scenarios involving several participants engaged in formal and informal dialogue, is introduced. The proposed method is thoroughly tested and benchmarked with respect to several state-of-the art diarization algorithms.
Israel D. Gebru, Sileye O. Ba, Xiaofei Li 0001, Radu Horaud
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 EM Algorithms for Weighted-Data Clustering with Application to Audio-Visual Scene Analysis
abstract
Data clustering has received a lot of attention and numerous methods, algorithms and software packages are available. Among these techniques, parametric finite-mixture models play a central role due to their interesting mathematical properties and to the existence of maximum-likelihood estimators based on expectation-maximization (EM). In this paper we propose a new mixture model that associates a weight with each observed point. We introduce the weighted-data Gaussian mixture and we derive two EM algorithms. The first one considers a fixed weight for each observation. The second one treats each weight as a random variable following a gamma distribution. We propose a model selection method based on a minimum message length criterion, provide a weight initialization strategy, and validate the proposed algorithms by comparing them with several state of the art parametric and non-parametric clustering techniques. We also demonstrate the effectiveness and robustness of the proposed clustering technique in the presence of heterogeneous data, namely audio-visual scene analysis.
Israel D. Gebru, Xavier Alameda-Pineda, Florence Forbes, Radu Horaud
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 A Distributed Architecture for Interacting with NAO
abstract
One of the main applications of the humanoid robot NAO - a small robot companion - is human-robot interaction (HRI). NAO is particularly well suited for HRI applications because of its design, hardware specifications, programming capabilities, and affordable cost. Indeed, NAO can stand up, walk, wander, dance, play soccer, sit down, recognize and grasp simple objects, detect and identify people, localize sounds, understand some spoken words, engage itself in simple and goal-directed dialogs, and synthesize speech. This is made possible due to the robot's 24 degree-of-freedom articulated structure (body, legs, feet, arms, hands, head, etc.), motors, cameras, microphones, etc., as well as to its on-board computing hardware and embedded software, e.g., robot motion control. Nevertheless, the current NAO configuration has two drawbacks that restrict the complexity of interactive behaviors that could potentially be implemented. Firstly, the on-board computing resources are inherently limited, which implies that it is difficult to implement sophisticated computer vision and audio signal analysis algorithms required by advanced interactive tasks. Secondly, programming new robot functionalities currently implies the development of embedded software, which is a difficult task in its own right necessitating specialized knowledge. The vast majority of HRI practitioners may not have this kind of expertise and hence they cannot easily and quickly implement their ideas, carry out thorough experimental validations, and design proof-of-concept demonstrators. We have developed a distributed software architecture that attempts to overcome these two limitations. Broadly speaking, NAO's on-board computing resources are augmented with external computing resources. The latter is a computer platform with its CPUs, GPUs, memory, operating system, libraries, software packages, internet access, etc. This configuration enables easy and fast development in Matlab, C, C++, or Python. Moreover, it allows the user to combine on-board libraries (motion control, face detection, etc.) with external toolboxes, e.g., OpenCv.
Fabien Badeig, Quentin Pelorson, Soraya Arias, Vincent Drouard, Israel D. Gebru, Xiaofei Li 0001, Georgios Evangelidis 0002, Radu Horaud
ICMI5
2013 Counter-forensics of median filtering
abstract
Median filtering is a well-known non linear denoising filter often used as an harmless post-processing, sometimes also employed to affect the reliability of some forensic techniques. In this work, we present a novel counter-forensic method able to conceal the characteristic traces left by median filtering. By exploiting the knowledge of features used in existing median filtering detectors, we are able to remove the characteristic footprints via suitable random pixel modification, while keeping the quality of the counter-attacked image high. Experimental results show that the proposed method is very effective, computationally efficient and competitive with other state-of-the-art techniques.
Duc-Tien Dang-Nguyen, Israel D. Gebru, Valentina Conotter, Giulia Boato, Francesco G. B. De Natale
MMSP2