EDBT 2026 Demo / reviewers in the wild / expert
Vladimir Tourbabin
dblp:119/7170
· DBLP profile ↗
16ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-2536-5666ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
8 papers |
Audio and music processing · 100% | |
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
speech enhancement |
2.6 | 4 | 2024 | Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
Audio and music processing
spatial audio |
1.8 | 3 | 2024 | Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Mixed Source Sound Field Translation for Virtual Binaural Application With Perceptual Validation · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Audio and music processing › spatial audio
binaural reproduction |
1.3 | 2 | 2024 | Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Mixed Source Sound Field Translation for Virtual Binaural Application With Perceptual Validation · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Audio and music processing › sound source localization
direction-of-arrival estimation |
1.0 | 2 | 2024 | Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information · IEEE ACM Trans. Audio Speech Lang. Process. 2024 Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots · IEEE ACM Trans. Audio Speech Lang. Process. 2015 |
Audio and music processing › speech enhancement › noise reduction
multichannel noise reduction |
0.8 | 1 | 2024 | Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Audio and music processing › speech enhancement
noise reduction |
0.8 | 1 | 2024 | Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Audio and music processing
beamforming |
0.7 | 2 | 2022 | Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays · IEEE Trans. Speech Audio Process. 2012 |
Audio and music processing
microphone array processing |
0.6 | 3 | 2015 | Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots · IEEE ACM Trans. Audio Speech Lang. Process. 2015 Theoretical framework for the optimization of microphone array configuration for humanoid robot audition · IEEE ACM Trans. Audio Speech Lang. Process. 2014 Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays · IEEE Trans. Speech Audio Process. 2012 |
Audio and music processing › beamforming
acoustic beamforming |
0.5 | 1 | 2021 | Robustness of Acoustic Rake Filters in Minimum Variance Beamforming · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
reverberation suppression |
0.5 | 1 | 2021 | Robustness of Acoustic Rake Filters in Minimum Variance Beamforming · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Audio and music processing › microphone array processing
array geometry optimization |
0.2 | 1 | 2014 | Theoretical framework for the optimization of microphone array configuration for humanoid robot audition · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition |
0.1 | 2 | 2015 | Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots · IEEE ACM Trans. Audio Speech Lang. Process. 2015 Theoretical framework for the optimization of microphone array configuration for humanoid robot audition · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Methods — techniques the papers use, named apart from their topics
short-time fourier transform · 0.8neural network · 0.8multichannel wiener filter · 0.8model-based deep learning · 0.8frequency smoothing · 0.6acoustic transfer function estimation · 0.6near-field and far-field source mixture · 0.5monte carlo simulation · 0.5blind parameter estimation · 0.5MUSHRA perceptual validation · 0.5signal model · 0.2motion compensation · 0.2generalized head related transfer functions · 0.2effective rank · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Importance of Spatial and Spectral Information in Multiple Speaker TrackingabstractMulti-speaker localization and tracking using microphone array recording is of importance in a wide range of applications. One of the challenges with multi-speaker tracking is to associate direction estimates with the correct speaker. Most existing association approaches rely on spatial or spectral information alone, leading to performance degradation when one of these information channels is partially known or missing. This paper studies a joint probability data association (JPDA)-based method that facilitates association based on joint spatial-spectral information. This is achieved by integrating speaker time-frequency (TF) masks, estimated based on spectral information, in the association probabilities calculation. An experimental study that tested the proposed method on recordings from the LOCATA challenge demonstrates the enhanced performance obtained by using joint spatial-spectral information in the association. Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely |
ICASSP | 2 |
| 2024 | Ambisonics Networks - The Effect of Radial Functions RegularizationabstractAmbisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The process of encoding Ambisonics from a spherical microphone array involves dividing by the radial functions, which may amplify noise at low frequencies. This can be overcome by regularization, with the downside of introducing errors to the Ambisonics encoding. This paper aims to investigate the impact of different ways of regularization on Deep Neural Network (DNN) training and performance. Ideally, these networks should be robust to the way of regularization. Simulated data of a single speaker in a room and experimental data from the LOCATA challenge were used to evaluate this robustness on an example algorithm of speaker localization based on the direct-path dominance (DPD) test. Results show that performance may be sensitive to the way of regularization, and an informed approach is proposed and investigated, highlighting the importance of regularization information. Bar Shaybet, Anurag Kumar 0003, Vladimir Tourbabin, Boaz Rafaely |
ICASSP | 3 |
| 2024 | Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural ReproductionabstractReal-life acoustic scenes may be recorded with microphone arrays for spatial audio applications, especially for the purpose of reproducing binaural signals for headphone listening. However, the presence of noise and interference may necessitate preprocessing to enhance the desired signal and improve the listener experience. Various methods have been developed to reduce noise while preserving the desired signal component with minimal distortion. The additional challenges posed by time-varying acoustic scenes are commonly addressed by segmenting the recorded signals into short time frames. Then, the short-time Fourier transform (STFT) is employed with multi-channel Wiener filter (MWF) and assuming the multiplicative transfer function (MTF) approximation. This approximation may not apply in the presence of long reverberation times and/or short STFT frames, so alternative techniques are required. This paper explores MWF-based enhancement in time-varying acoustic scenes where the MTF approximation is inapplicable, both analytically and experimentally with normal-hearing listeners. The investigated scene comprises a single desired source in a reverberant environment, and the impact of frame length and acoustic parameters on the rank of the spatial covariance matrix is studied. It is revealed that superior results in terms of reduced distortion and improved listener experience are achieved when using a full-rank spatial covariance matrix. Moti Lugasi, Jacob Donley, Anjali Menon, Vladimir Tourbabin, Boaz Rafaely |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial InformationabstractDirection-of-arrival (DOA) estimation is a fundamental task in audio signal processing that becomes difficult in real-world environments due to the presence of reverberation. To address this difficulty, Direct-Path Dominance (DPD) tests have been proposed as an effective approach for detecting time-frequency (TF) bins dominated by direct sound, which contain accurate DOA information. These have been found to be particularly efficient when working with spherical arrays. While methods based on neural networks (NNs) have been developed to estimate the DOA, they have limitations such as the need for a large training database, and often understanding of the system's operation is lacking. This work proposes two novel DPD-test methods based on a model-based deep learning approach that combines the original DPD-test model with a data-driven system. Thus, it is possible to preserve the robustness of the original DPD-test across acoustic environments, while using a data-driven approach to better extract useful information about the direct sound, thereby enhancing the original method's performance. In particular, the paper investigates how energetic, temporal and spatial information contribute to the identification of TF-bins dominated by the direct signal. The proposed methods are trained on simulated data of a single sound source in a room, and evaluated on simulated and real data. The results show that energetic and temporal information provide new information about direct sound, which has not been considered in previous works and can improve its performance. Orel Ben Zaken, Anurag Kumar 0003, Vladimir Tourbabin, Boaz Rafaely |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Subspace Hybrid Beamforming for Head-Worn Microphone ArraysabstractA two-stage multi-channel speech enhancement method is proposed which consists of a novel adaptive beamformer, Hybrid Minimum Variance Distortionless Response (MVDR), Isotropic-MVDR (Iso), and a novel multi-channel spectral Principal Components Analysis (PCA) denoising. In the first stage, the Hybrid-MVDR performs multiple MVDRs using a dictionary of pre-defined noise field models and picks the minimum-power outcome, which benefits from the robustness of signal-independent beamforming and the performance of adaptive beamforming. In the second stage, the outcomes of Hybrid and Iso are jointly used in a two-channel PCA-based denoising to remove the ‘musical noise’ produced by Hybrid beamformer. On a dataset of real ‘cocktail-party’ recordings with head-worn array, the proposed method outperforms the baseline superdirective beamformer in noise suppression (fwSegSNR, SDR, SIR, SAR) and speech intelligibility (STOI) with similar speech quality (PESQ) improvement. Sina Hafezi, Alastair H. Moore, Pierre Guiraud, Patrick A. Naylor, Jacob Donley, Vladimir Tourbabin, Thomas Lunner |
ICASSP | 6 |
| 2022 | Audio Signal Processing for Telepresence Based on Wearable Array in Noisy and Dynamic ScenesabstractTelepresence for virtual meetings has gained interest due to recent travel limitations and the new reality of working from home. However, current literature supporting real-world microphone arrays for realistic telepresence in audio is very limited. This paper investigates a scenario of a distant participant joining virtually a meeting between two dynamic participants. The audio signal processing chain (i) starts by recording using an array mounted on glasses, (ii) with initial processing providing direction-of-arrival estimation of a desired speaker using a direct-path dominance test robust to reverberation, combined with speaker separation for improved dynamic localization, (iii) followed by speech enhancement against interfering speakers and noise, (iv) and ends with applying binaural signal matching for headphone listening. This paper compares model-based processing to learning-based processing in both noisy and dynamic scenarios, and presents a novel processing using data from a real wearable array, studied by simulation and a listening test. Hanan Beit-On, Moti Lugasi, Lior Madmoni, Anjali Menon, Anurag Kumar 0003, Jacob Donley, Vladimir Tourbabin, Boaz Rafaely |
ICASSP | 7 |
| 2022 | Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved PerceptionabstractIn many applications, such as hearing aids and virtual reality, spatial audio is used to provide a more natural experience to the users. However, when captured in the real world, the audio signals may suffer from noise and interference. In this case, the challenge is to attenuate the undesired signals, while preserving the desired signals with their spatial information. In this paper, an approach for spatial signal enhancement is presented. This approach is based on two phases of estimation. The first phase is source signal estimation using a beamformer. Then, in the second phase, the acoustic transfer function (ATF) between the source and the array is estimated; this leads to an enhanced estimation of the desired signal at the microphones. This approach has been previously proposed, but was not investigated in depth. In this paper, a model for the estimated desired signals is developed. In contrast to other methods of spatial enhancement, no trade-off between noise reduction and signal distortion is found in this model, when there is a single desired source and a single interfering source in a reverberant room. To overcome the limited accuracy of ATF estimation for short duration signals, frequency smoothing is applied. Listening tests verify the performance of the proposed approach. Moti Lugasi, Anjali Menon, Vladimir Tourbabin, Boaz Rafaely |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Mixed Source Sound Field Translation for Virtual Binaural Application With Perceptual ValidationabstractNon-interactive and linear experienceslike cinema film offer high quality surround sound audio to enhance immersion, however, the perspective is usually fixed to the recording microphone position. With the rise of virtual reality, there is a demand for recording and recreating real-world experiences that allow users to move throughout the reproduction. Sound field translation achieves this by building an equivalent environment of virtual sources to recreate the recording spatially. However, the technique remains to restrict the maximum distance a user can translate away from the recording microphone's perspective due to the discrete sampling by commercial higher order microphones only being capable of recording an acoustic sweet-spot. In this paper, we propose a method for binaurally reproducing a microphone recording in a virtual application that allows the user to freely translate their body further beyond the recording position. The method incorporates a mixture of near-field and far-field sources in a sparsely expanded virtual environment to maintain a perceptually accurate reproduction. We perceptually validate the method through a Multiple Stimulus with Hidden Reference and Anchor (MUSHRA) experiment. Compared to the planewave benchmark, the proposed method offers both improved source localizability and robustness to spectral distortions at translated listening positions. A cross-examination with numerical simulations demonstrated that the sparse expansion relaxes the inherent sweet-spot constraint, leading to the improved localizability for sparse environments. Additionally, the proposed method is seen to better reproduce the intensity and binaural room impulse response spectra of near-field environments, further supporting the perceptual results. Lachlan Birnie, Thushara D. Abhayapala, Vladimir Tourbabin, Prasanga N. Samarasinghe |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Robustness of Acoustic Rake Filters in Minimum Variance BeamformingabstractAcoustic rake filters perform coherent summation of the early room reflections using beamforming, with the aim of improving beamforming performance. This concept has been investigated for speech enhancement applications, improving noise reduction and late reverberation attenuation. Current studies typically assume that the parameters of the early reflections, such as the direction-of-arrival, delay and amplitude, are known in advance in the rake filter design. This work presents a novel investigation of the acoustic rake filter in a more practical context, focusing on the minimum variance distortionless response (MVDR) formulation. First, the sensitivity of the filter performance to perturbations in the reflection parameters is derived analytically, and investigated numerically using Monte Carlo simulations with a spherical microphone array. Then, an end-to-end example of rake filtering in a blind scenario is presented, where the reflection parameters are estimated from speech signals without any prior information. This example demonstrates for the first time the use of rake filtering in a realistic scenario. Ran Weisman, Tom Shlomo, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | The Importance of Time-Frequency Averaging for Binaural Speaker Localization in Reverberant Environments
Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely |
INTERSPEECH | 2 |
| 2020 | Spatial Covariance Matrix Estimation for Reverberant Speech with Application to Speech Enhancement
Ran Weisman, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely |
INTERSPEECH | 2 |
| 2015 | Enhanced robot audition by dynamic acoustic sensing in moving humanoidsabstractAuditory systems of humanoid robots usually acquire the surrounding sound field by means of microphone arrays. These arrays can undergo motion related to the robot's activity. The conventional approach to dealing with this motion is to stop the robot during sound acquisition. This approach avoids changing the positions of the microphones during the acquisition and reduces the robot's ego-noise. However, stopping the robot can interfere with the naturalness of its behaviour. Moreover, the potential performance improvement due to motion of the sound acquiring system can not be attained. This potential is analysed in the current paper. The analysis considers two different types of motion: (i) rotation of the robot's head and (ii) limb gestures. The study presented here combines both theoretical and numerical simulation approaches. The results show that rotation of the head improves the high-frequency performance of the microphone array positioned on the head of the robot. This is complemented by the limb gestures, which improve the low-frequency performance of the array positioned on the torso and limbs of the robot. Vladimir Tourbabin, Hendrik Barfuss, Boaz Rafaely, Walter Kellermann |
ICASSP | 1 |
| 2015 | Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid RobotsabstractThe auditory system of humanoid robots has gained increased attention in recent years. This system typically acquires the surrounding sound field by means of a microphone array. Signals acquired by the array are then processed using various methods. One of the widely applied methods is direction of arrival estimation. The conventional direction of arrival estimation methods assume that the array is fixed at a given position during the estimation. However, this is not necessarily true for an array installed on a moving humanoid robot. The array motion, if not accounted for appropriately, can introduce a significant error in the estimated direction of arrival. The current paper presents a signal model that takes the motion into account. Based on this model, two processing methods are proposed. The first one compensates for the motion of the robot. The second method is applicable to periodic signals and utilizes the motion in order to enhance the performance to a level beyond that of a stationary array. Numerical simulations and an experimental study are provided, demonstrating that the motion compensation method almost eliminates the motion-related error. It is also demonstrated that by using the motion-based enhancement method it is possible to improve the direction of arrival estimation performance, as compared to that obtained when using a stationary array. Vladimir Tourbabin, Boaz Rafaely |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Theoretical framework for the optimization of microphone array configuration for humanoid robot auditionabstractAn important aspect of a humanoid robot is audition. Previous work has presented robot systems capable of sound localization and source segregation based on microphone arrays with various configurations. However, no theoretical framework for the design of these arrays has been presented. In the current paper, a design framework is proposed based on a novel array quality measure. The measure is based on the effective rank of a matrix composed of the generalized head related transfer functions (GHRTFs) that account for microphone positions other than the ears. The measure is shown to be theoretically related to standard array performance measures such as beamforming robustness and DOA estimation accuracy. Then, the measure is applied to produce sample designs of microphone arrays. Their performance is investigated numerically, verifying the advantages of array design based on the proposed theoretical framework. Vladimir Tourbabin, Boaz Rafaely |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Theoretical framework for the design of microphone arrays for robot auditionabstractAn important part of a human-like robot is robot audition. Previous work presented systems capable of sound localization and source segregation based on microphone arrays of various configurations. However, no theoretical framework for assessing the quality of these array configurations has been presented. In the current paper such a measure is proposed based on the generalized HRTFs that account for microphone positions other than the ears. The measure is analyzed theoretically with respect to beamforming robustness and DOA estimation accuracy. The measure is then used to find the optimal location of a single microphone and a pair of microphones based on the generalized HRTF database obtained by means of BEM simulation. The results are not surprising, showing that the best position of a single microphone is the ear canal. For a pair of microphones, the results generally show that the sensors should be maximally spatially separated. Vladimir Tourbabin, Boaz Rafaely |
ICASSP | 1 |
| 2012 | Optimal Real-Weighted Beamforming With Application to Linear and Spherical ArraysabstractOne of the uses of sensor arrays is for spatial filtering or beamforming. Current digital signal processing methods facilitate complex-weighted beamforming, providing flexibility in array design. Previous studies proposed the use of real-valued beamforming weights, which although reduce flexibility in design, may provide a range of benefits, e.g., simplified beamformer implementation or efficient beamforming algorithms. This paper presents a new method for the design of arrays with real-valued weights, that achieve maximum directivity, providing closed-form solution to array weights. The method is studied for linear and spherical arrays, where it is shown that rigid spherical arrays are particularly suitable for real-weight designs as they do not suffer from grating lobes, a dominant feature in linear arrays with real weights. A simulation study is presented for linear and spherical arrays, along with an experimental investigation, validating the theoretical developments. Vladimir Tourbabin, Morag Agmon, Boaz Rafaely, Joseph Tabrikian |
IEEE Trans. Speech Audio Process. | 1 |