Vladimir Tourbabin

dblp:119/7170 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-2536-5666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
8 papers
Audio and music processing · 100%
Artificial intelligence
2 papers
Speech recognition and synthesis · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech enhancement
2.642024
Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Audio and music processing
spatial audio
1.832024
Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Mixed Source Sound Field Translation for Virtual Binaural Application With Perceptual Validation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing › spatial audio
binaural reproduction
1.322024
Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Mixed Source Sound Field Translation for Virtual Binaural Application With Perceptual Validation · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing › sound source localization
direction-of-arrival estimation
1.022024
Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Audio and music processing › speech enhancement › noise reduction
multichannel noise reduction
0.812024
Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Audio and music processing › speech enhancement
noise reduction
0.812024
Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Audio and music processing
beamforming
0.722022
Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays · IEEE Trans. Speech Audio Process. 2012
Audio and music processing
microphone array processing
0.632015
Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Theoretical framework for the optimization of microphone array configuration for humanoid robot audition · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays · IEEE Trans. Speech Audio Process. 2012
Audio and music processing › beamforming
acoustic beamforming
0.512021
Robustness of Acoustic Rake Filters in Minimum Variance Beamforming · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing › acoustic signal processing › audio signal reconstruction › audio restoration
reverberation suppression
0.512021
Robustness of Acoustic Rake Filters in Minimum Variance Beamforming · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Audio and music processing › microphone array processing
array geometry optimization
0.212014
Theoretical framework for the optimization of microphone array configuration for humanoid robot audition · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Natural language and speech › Speech recognition and synthesis › speech separation › computational auditory scene analysis
robot audition
0.122015
Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Theoretical framework for the optimization of microphone array configuration for humanoid robot audition · IEEE ACM Trans. Audio Speech Lang. Process. 2014

Methods — techniques the papers use, named apart from their topics

short-time fourier transform · 0.8neural network · 0.8multichannel wiener filter · 0.8model-based deep learning · 0.8frequency smoothing · 0.6acoustic transfer function estimation · 0.6near-field and far-field source mixture · 0.5monte carlo simulation · 0.5blind parameter estimation · 0.5MUSHRA perceptual validation · 0.5signal model · 0.2motion compensation · 0.2generalized head related transfer functions · 0.2effective rank · 0.2
YearPublicationVenuePosition
2025 The Importance of Spatial and Spectral Information in Multiple Speaker Tracking
abstract
Multi-speaker localization and tracking using microphone array recording is of importance in a wide range of applications. One of the challenges with multi-speaker tracking is to associate direction estimates with the correct speaker. Most existing association approaches rely on spatial or spectral information alone, leading to performance degradation when one of these information channels is partially known or missing. This paper studies a joint probability data association (JPDA)-based method that facilitates association based on joint spatial-spectral information. This is achieved by integrating speaker time-frequency (TF) masks, estimated based on spectral information, in the association probabilities calculation. An experimental study that tested the proposed method on recordings from the LOCATA challenge demonstrates the enhanced performance obtained by using joint spatial-spectral information in the association.
Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
ICASSP2
2024 Ambisonics Networks - The Effect of Radial Functions Regularization
abstract
Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The process of encoding Ambisonics from a spherical microphone array involves dividing by the radial functions, which may amplify noise at low frequencies. This can be overcome by regularization, with the downside of introducing errors to the Ambisonics encoding. This paper aims to investigate the impact of different ways of regularization on Deep Neural Network (DNN) training and performance. Ideally, these networks should be robust to the way of regularization. Simulated data of a single speaker in a room and experimental data from the LOCATA challenge were used to evaluate this robustness on an example algorithm of speaker localization based on the direct-path dominance (DPD) test. Results show that performance may be sensitive to the way of regularization, and an informed approach is proposed and investigated, highlighting the importance of regularization information.
Bar Shaybet, Anurag Kumar 0003, Vladimir Tourbabin, Boaz Rafaely
ICASSP3
2024 Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction
abstract
Real-life acoustic scenes may be recorded with microphone arrays for spatial audio applications, especially for the purpose of reproducing binaural signals for headphone listening. However, the presence of noise and interference may necessitate preprocessing to enhance the desired signal and improve the listener experience. Various methods have been developed to reduce noise while preserving the desired signal component with minimal distortion. The additional challenges posed by time-varying acoustic scenes are commonly addressed by segmenting the recorded signals into short time frames. Then, the short-time Fourier transform (STFT) is employed with multi-channel Wiener filter (MWF) and assuming the multiplicative transfer function (MTF) approximation. This approximation may not apply in the presence of long reverberation times and/or short STFT frames, so alternative techniques are required. This paper explores MWF-based enhancement in time-varying acoustic scenes where the MTF approximation is inapplicable, both analytically and experimentally with normal-hearing listeners. The investigated scene comprises a single desired source in a reverberant environment, and the impact of frame length and acoustic parameters on the rank of the spatial covariance matrix is studied. It is revealed that superior results in terms of reduced distortion and improved listener experience are achieved when using a full-rank spatial covariance matrix.
Moti Lugasi, Jacob Donley, Anjali Menon, Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information
abstract
Direction-of-arrival (DOA) estimation is a fundamental task in audio signal processing that becomes difficult in real-world environments due to the presence of reverberation. To address this difficulty, Direct-Path Dominance (DPD) tests have been proposed as an effective approach for detecting time-frequency (TF) bins dominated by direct sound, which contain accurate DOA information. These have been found to be particularly efficient when working with spherical arrays. While methods based on neural networks (NNs) have been developed to estimate the DOA, they have limitations such as the need for a large training database, and often understanding of the system's operation is lacking. This work proposes two novel DPD-test methods based on a model-based deep learning approach that combines the original DPD-test model with a data-driven system. Thus, it is possible to preserve the robustness of the original DPD-test across acoustic environments, while using a data-driven approach to better extract useful information about the direct sound, thereby enhancing the original method's performance. In particular, the paper investigates how energetic, temporal and spatial information contribute to the identification of TF-bins dominated by the direct signal. The proposed methods are trained on simulated data of a single sound source in a room, and evaluated on simulated and real data. The results show that energetic and temporal information provide new information about direct sound, which has not been considered in previous works and can improve its performance.
Orel Ben Zaken, Anurag Kumar 0003, Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Subspace Hybrid Beamforming for Head-Worn Microphone Arrays
abstract
A two-stage multi-channel speech enhancement method is proposed which consists of a novel adaptive beamformer, Hybrid Minimum Variance Distortionless Response (MVDR), Isotropic-MVDR (Iso), and a novel multi-channel spectral Principal Components Analysis (PCA) denoising. In the first stage, the Hybrid-MVDR performs multiple MVDRs using a dictionary of pre-defined noise field models and picks the minimum-power outcome, which benefits from the robustness of signal-independent beamforming and the performance of adaptive beamforming. In the second stage, the outcomes of Hybrid and Iso are jointly used in a two-channel PCA-based denoising to remove the ‘musical noise’ produced by Hybrid beamformer. On a dataset of real ‘cocktail-party’ recordings with head-worn array, the proposed method outperforms the baseline superdirective beamformer in noise suppression (fwSegSNR, SDR, SIR, SAR) and speech intelligibility (STOI) with similar speech quality (PESQ) improvement.
Sina Hafezi, Alastair H. Moore, Pierre Guiraud, Patrick A. Naylor, Jacob Donley, Vladimir Tourbabin, Thomas Lunner
ICASSP6
2022 Audio Signal Processing for Telepresence Based on Wearable Array in Noisy and Dynamic Scenes
abstract
Telepresence for virtual meetings has gained interest due to recent travel limitations and the new reality of working from home. However, current literature supporting real-world microphone arrays for realistic telepresence in audio is very limited. This paper investigates a scenario of a distant participant joining virtually a meeting between two dynamic participants. The audio signal processing chain (i) starts by recording using an array mounted on glasses, (ii) with initial processing providing direction-of-arrival estimation of a desired speaker using a direct-path dominance test robust to reverberation, combined with speaker separation for improved dynamic localization, (iii) followed by speech enhancement against interfering speakers and noise, (iv) and ends with applying binaural signal matching for headphone listening. This paper compares model-based processing to learning-based processing in both noisy and dynamic scenarios, and presents a novel processing using data from a real wearable array, studied by simulation and a listening test.
Hanan Beit-On, Moti Lugasi, Lior Madmoni, Anjali Menon, Anurag Kumar 0003, Jacob Donley, Vladimir Tourbabin, Boaz Rafaely
ICASSP7
2022 Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception
abstract
In many applications, such as hearing aids and virtual reality, spatial audio is used to provide a more natural experience to the users. However, when captured in the real world, the audio signals may suffer from noise and interference. In this case, the challenge is to attenuate the undesired signals, while preserving the desired signals with their spatial information. In this paper, an approach for spatial signal enhancement is presented. This approach is based on two phases of estimation. The first phase is source signal estimation using a beamformer. Then, in the second phase, the acoustic transfer function (ATF) between the source and the array is estimated; this leads to an enhanced estimation of the desired signal at the microphones. This approach has been previously proposed, but was not investigated in depth. In this paper, a model for the estimated desired signals is developed. In contrast to other methods of spatial enhancement, no trade-off between noise reduction and signal distortion is found in this model, when there is a single desired source and a single interfering source in a reverberant room. To overcome the limited accuracy of ATF estimation for short duration signals, frequency smoothing is applied. Listening tests verify the performance of the proposed approach.
Moti Lugasi, Anjali Menon, Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Mixed Source Sound Field Translation for Virtual Binaural Application With Perceptual Validation
abstract
Non-interactive and linear experienceslike cinema film offer high quality surround sound audio to enhance immersion, however, the perspective is usually fixed to the recording microphone position. With the rise of virtual reality, there is a demand for recording and recreating real-world experiences that allow users to move throughout the reproduction. Sound field translation achieves this by building an equivalent environment of virtual sources to recreate the recording spatially. However, the technique remains to restrict the maximum distance a user can translate away from the recording microphone's perspective due to the discrete sampling by commercial higher order microphones only being capable of recording an acoustic sweet-spot. In this paper, we propose a method for binaurally reproducing a microphone recording in a virtual application that allows the user to freely translate their body further beyond the recording position. The method incorporates a mixture of near-field and far-field sources in a sparsely expanded virtual environment to maintain a perceptually accurate reproduction. We perceptually validate the method through a Multiple Stimulus with Hidden Reference and Anchor (MUSHRA) experiment. Compared to the planewave benchmark, the proposed method offers both improved source localizability and robustness to spectral distortions at translated listening positions. A cross-examination with numerical simulations demonstrated that the sparse expansion relaxes the inherent sweet-spot constraint, leading to the improved localizability for sparse environments. Additionally, the proposed method is seen to better reproduce the intensity and binaural room impulse response spectra of near-field environments, further supporting the perceptual results.
Lachlan Birnie, Thushara D. Abhayapala, Vladimir Tourbabin, Prasanga N. Samarasinghe
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Robustness of Acoustic Rake Filters in Minimum Variance Beamforming
abstract
Acoustic rake filters perform coherent summation of the early room reflections using beamforming, with the aim of improving beamforming performance. This concept has been investigated for speech enhancement applications, improving noise reduction and late reverberation attenuation. Current studies typically assume that the parameters of the early reflections, such as the direction-of-arrival, delay and amplitude, are known in advance in the rake filter design. This work presents a novel investigation of the acoustic rake filter in a more practical context, focusing on the minimum variance distortionless response (MVDR) formulation. First, the sensitivity of the filter performance to perturbations in the reflection parameters is derived analytically, and investigated numerically using Monte Carlo simulations with a spherical microphone array. Then, an end-to-end example of rake filtering in a blind scenario is presented, where the reflection parameters are estimated from speech signals without any prior information. This example demonstrates for the first time the use of rake filtering in a realistic scenario.
Ran Weisman, Tom Shlomo, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 The Importance of Time-Frequency Averaging for Binaural Speaker Localization in Reverberant Environments
Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
INTERSPEECH2
2020 Spatial Covariance Matrix Estimation for Reverberant Speech with Application to Speech Enhancement
Ran Weisman, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely
INTERSPEECH2
2015 Enhanced robot audition by dynamic acoustic sensing in moving humanoids
abstract
Auditory systems of humanoid robots usually acquire the surrounding sound field by means of microphone arrays. These arrays can undergo motion related to the robot's activity. The conventional approach to dealing with this motion is to stop the robot during sound acquisition. This approach avoids changing the positions of the microphones during the acquisition and reduces the robot's ego-noise. However, stopping the robot can interfere with the naturalness of its behaviour. Moreover, the potential performance improvement due to motion of the sound acquiring system can not be attained. This potential is analysed in the current paper. The analysis considers two different types of motion: (i) rotation of the robot's head and (ii) limb gestures. The study presented here combines both theoretical and numerical simulation approaches. The results show that rotation of the head improves the high-frequency performance of the microphone array positioned on the head of the robot. This is complemented by the limb gestures, which improve the low-frequency performance of the array positioned on the torso and limbs of the robot.
Vladimir Tourbabin, Hendrik Barfuss, Boaz Rafaely, Walter Kellermann
ICASSP1
2015 Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots
abstract
The auditory system of humanoid robots has gained increased attention in recent years. This system typically acquires the surrounding sound field by means of a microphone array. Signals acquired by the array are then processed using various methods. One of the widely applied methods is direction of arrival estimation. The conventional direction of arrival estimation methods assume that the array is fixed at a given position during the estimation. However, this is not necessarily true for an array installed on a moving humanoid robot. The array motion, if not accounted for appropriately, can introduce a significant error in the estimated direction of arrival. The current paper presents a signal model that takes the motion into account. Based on this model, two processing methods are proposed. The first one compensates for the motion of the robot. The second method is applicable to periodic signals and utilizes the motion in order to enhance the performance to a level beyond that of a stationary array. Numerical simulations and an experimental study are provided, demonstrating that the motion compensation method almost eliminates the motion-related error. It is also demonstrated that by using the motion-based enhancement method it is possible to improve the direction of arrival estimation performance, as compared to that obtained when using a stationary array.
Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Theoretical framework for the optimization of microphone array configuration for humanoid robot audition
abstract
An important aspect of a humanoid robot is audition. Previous work has presented robot systems capable of sound localization and source segregation based on microphone arrays with various configurations. However, no theoretical framework for the design of these arrays has been presented. In the current paper, a design framework is proposed based on a novel array quality measure. The measure is based on the effective rank of a matrix composed of the generalized head related transfer functions (GHRTFs) that account for microphone positions other than the ears. The measure is shown to be theoretically related to standard array performance measures such as beamforming robustness and DOA estimation accuracy. Then, the measure is applied to produce sample designs of microphone arrays. Their performance is investigated numerically, verifying the advantages of array design based on the proposed theoretical framework.
Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Theoretical framework for the design of microphone arrays for robot audition
abstract
An important part of a human-like robot is robot audition. Previous work presented systems capable of sound localization and source segregation based on microphone arrays of various configurations. However, no theoretical framework for assessing the quality of these array configurations has been presented. In the current paper such a measure is proposed based on the generalized HRTFs that account for microphone positions other than the ears. The measure is analyzed theoretically with respect to beamforming robustness and DOA estimation accuracy. The measure is then used to find the optimal location of a single microphone and a pair of microphones based on the generalized HRTF database obtained by means of BEM simulation. The results are not surprising, showing that the best position of a single microphone is the ear canal. For a pair of microphones, the results generally show that the sensors should be maximally spatially separated.
Vladimir Tourbabin, Boaz Rafaely
ICASSP1
2012 Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays
abstract
One of the uses of sensor arrays is for spatial filtering or beamforming. Current digital signal processing methods facilitate complex-weighted beamforming, providing flexibility in array design. Previous studies proposed the use of real-valued beamforming weights, which although reduce flexibility in design, may provide a range of benefits, e.g., simplified beamformer implementation or efficient beamforming algorithms. This paper presents a new method for the design of arrays with real-valued weights, that achieve maximum directivity, providing closed-form solution to array weights. The method is studied for linear and spherical arrays, where it is shown that rigid spherical arrays are particularly suitable for real-weight designs as they do not suffer from grating lobes, a dominant feature in linear arrays with real weights. A simulation study is presented for linear and spherical arrays, along with an experimental investigation, validating the theoretical developments.
Vladimir Tourbabin, Morag Agmon, Boaz Rafaely, Joseph Tabrikian
IEEE Trans. Speech Audio Process.1