Boaz Rafaely

dblp:42/3482 · DBLP profile ↗
← Back
57ranked-venue papers
12as first author
15since 2021 · last 2025
0000-0002-6819-9250ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 7 since 2021
YearPublicationVenuePosition
2025 The Importance of Spatial and Spectral Information in Multiple Speaker Tracking
abstract
Multi-speaker localization and tracking using microphone array recording is of importance in a wide range of applications. One of the challenges with multi-speaker tracking is to associate direction estimates with the correct speaker. Most existing association approaches rely on spatial or spectral information alone, leading to performance degradation when one of these information channels is partially known or missing. This paper studies a joint probability data association (JPDA)-based method that facilitates association based on joint spatial-spectral information. This is achieved by integrating speaker time-frequency (TF) masks, estimated based on spectral information, in the association probabilities calculation. An experimental study that tested the proposed method on recordings from the LOCATA challenge demonstrates the enhanced performance obtained by using joint spatial-spectral information in the association.
Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
ICASSP3
2025 Ambisonics Binaural Rendering via Masked Magnitude Least Squares
abstract
Ambisonics rendering has become an integral part of 3D audio for headphones. It works well with existing recording hardware, the processing cost is mostly independent of the number of sound sources, and it elegantly allows for rotating the scene and listener. One challenge in Ambisonics headphone rendering is to find a perceptually well behaved low-order representation of the Head-Related Transfer Functions (HRTFs) that are contained in the rendering pipe-line. Low-order rendering is of interest, when working with microphone arrays containing only a few sensors, or for reducing the bandwidth for signal transmission. Magnitude Least Squares rendering became the de facto standard for this, which discards high-frequency interaural phase information in favor of reducing magnitude errors. Building upon this idea, we suggest Masked Magnitude Least Squares, which optimized the Ambisonics coefficients with a neural network and employs a spatio-spectral weighting mask to control the accuracy of the magnitude reconstruction. In the tested case, the weighting mask helped to maintain high-frequency notches in the low-order HRTFs and improved the modeled median plane localization performance in comparison to MagLS, while only marginally affecting the overall accuracy of the magnitude reconstruction.
Or Berebi, Fabian Brinkmann, Stefan Weinzierl, Boaz Rafaely
ICASSP4
2024 On HRTF Notch Frequency Prediction using Anthropometric Features and Neural Networks
abstract
High fidelity spatial audio often performs better when produced using a personalized head-related transfer function (HRTF). However, the direct acquisition of HRTFs is cumbersome and requires specialized equipment. Thus, many personalization methods estimate HRTF features from easily obtained anthropometric features of the pinna, head, and torso. The first HRTF notch frequency (N1) is known to be a dominant feature in elevation localization, and thus a useful feature for HRTF personalization. This paper describes the prediction of N1 frequency from pinna anthropometry using a neural model. Prediction is performed separately on three databases, both simulated and measured, and then by domain mixing in-between the databases. The model successfully predicts N1 frequency for individual databases and by domain mixing between some databases. Prediction errors are better or comparable to those previously reported, showing significant improvement when acquired over a large database and with a larger output range.
Lior Arbel, Ishwarya Ananthabhotla, Zamir Ben-Hur, David L. Alon, Boaz Rafaely
ICASSP5
2024 Ambisonics Networks - The Effect of Radial Functions Regularization
abstract
Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The process of encoding Ambisonics from a spherical microphone array involves dividing by the radial functions, which may amplify noise at low frequencies. This can be overcome by regularization, with the downside of introducing errors to the Ambisonics encoding. This paper aims to investigate the impact of different ways of regularization on Deep Neural Network (DNN) training and performance. Ideally, these networks should be robust to the way of regularization. Simulated data of a single speaker in a room and experimental data from the LOCATA challenge were used to evaluate this robustness on an example algorithm of speaker localization based on the direct-path dominance (DPD) test. Results show that performance may be sensitive to the way of regularization, and an informed approach is proposed and investigated, highlighting the importance of regularization information.
Bar Shaybet, Anurag Kumar 0003, Vladimir Tourbabin, Boaz Rafaely
ICASSP4
2024 Analysis and Design of Head-Tracked Compensation for Bilateral Ambisonics
abstract
Virtual and augmented reality technologies demand high-quality spatial sound recording and playback through headphones. However, achieving high-quality binaural reproduction requires a complex recording system and a large number of microphones. To address this issue, a recent study proposed Bilateral Ambisonics, which involves capturing the sound-field using two low-order microphone arrays located ear-distance apart. We present an analytical analysis of the limitation of a previously suggested head-tracking compensation solution to Bilateral Ambisonics. An alternative approach is proposed to overcome these limits in which the translation operation is band-limited. A subjective evaluation and a listening test are provided and complement the findings of the analytical analysis. Results indicate that in a static scenario, compensating for small lateral head-rotations up to$\pm 30^{\circ }$with good accuracy is possible for microphone arrays of spherical harmonics (SH) order of 1 and for medium rotations of up to$\pm 60^{\circ }$with SH order of 2. When a dynamic scenario is considered, Bilateral Ambisonics of order 2 were comparable to High-order Ambisonics, and Bilateral Ambisonics of order 1 provided performance comparable to third order Ambisonics with MagLS.
Or Berebi, Zamir Ben-Hur, David L. Alon, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Multi-Channel to Multi-Channel Noise Reduction and Reverberant Speech Preservation in Time-Varying Acoustic Scenes for Binaural Reproduction
abstract
Real-life acoustic scenes may be recorded with microphone arrays for spatial audio applications, especially for the purpose of reproducing binaural signals for headphone listening. However, the presence of noise and interference may necessitate preprocessing to enhance the desired signal and improve the listener experience. Various methods have been developed to reduce noise while preserving the desired signal component with minimal distortion. The additional challenges posed by time-varying acoustic scenes are commonly addressed by segmenting the recorded signals into short time frames. Then, the short-time Fourier transform (STFT) is employed with multi-channel Wiener filter (MWF) and assuming the multiplicative transfer function (MTF) approximation. This approximation may not apply in the presence of long reverberation times and/or short STFT frames, so alternative techniques are required. This paper explores MWF-based enhancement in time-varying acoustic scenes where the MTF approximation is inapplicable, both analytically and experimentally with normal-hearing listeners. The investigated scene comprises a single desired source in a reverberant environment, and the impact of frame length and acoustic parameters on the rank of the spatial covariance matrix is studied. It is revealed that superior results in terms of reduced distortion and improved listener experience are achieved when using a full-rank spatial covariance matrix.
Moti Lugasi, Jacob Donley, Anjali Menon, Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 Neural-Network-Based Direction-of-Arrival Estimation for Reverberant Speech - The Importance of Energetic, Temporal, and Spatial Information
abstract
Direction-of-arrival (DOA) estimation is a fundamental task in audio signal processing that becomes difficult in real-world environments due to the presence of reverberation. To address this difficulty, Direct-Path Dominance (DPD) tests have been proposed as an effective approach for detecting time-frequency (TF) bins dominated by direct sound, which contain accurate DOA information. These have been found to be particularly efficient when working with spherical arrays. While methods based on neural networks (NNs) have been developed to estimate the DOA, they have limitations such as the need for a large training database, and often understanding of the system's operation is lacking. This work proposes two novel DPD-test methods based on a model-based deep learning approach that combines the original DPD-test model with a data-driven system. Thus, it is possible to preserve the robustness of the original DPD-test across acoustic environments, while using a data-driven approach to better extract useful information about the direct sound, thereby enhancing the original method's performance. In particular, the paper investigates how energetic, temporal and spatial information contribute to the identification of TF-bins dominated by the direct signal. The proposed methods are trained on simulated data of a single sound source in a room, and evaluated on simulated and real data. The results show that energetic and temporal information provide new information about direct sound, which has not been considered in previous works and can improve its performance.
Orel Ben Zaken, Anurag Kumar 0003, Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Weighted Frequency Smoothing for Enhanced Speaker Localization
abstract
The coherent signal subspace method may be used in order to apply subspace localization methods (e.g. MUSIC) to coherent sources. This method involves a focusing process followed by frequency smoothing, which is intended to decorrelate source signals from coherent sources. In practice, however, only moderate decorrelation is obtained, which may lead to performance degradation. Although decorrelation can be improved by widening the smoothing bandwidth, a wider bandwidth may increase focusing error and the smoothing bandwidth is limited by the bandwidth of the actual signal. In this paper, a weighted frequency smoothing that improves decorrelation for a given bandwidth is proposed. It is shown that better decorrelation is obtained by selecting the weights to be inversely proportional to the source signal power at the given frequency. However, since the power of the source is not known, it is estimated by the trace of the array spatial covariance matrix. An experimental study is presented that investigates the effect of the proposed weighting on DOA estimation of speech sources in a reverberant environment.
Hanan Beit-On, Tom Shlomo, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Audio Signal Processing for Telepresence Based on Wearable Array in Noisy and Dynamic Scenes
abstract
Telepresence for virtual meetings has gained interest due to recent travel limitations and the new reality of working from home. However, current literature supporting real-world microphone arrays for realistic telepresence in audio is very limited. This paper investigates a scenario of a distant participant joining virtually a meeting between two dynamic participants. The audio signal processing chain (i) starts by recording using an array mounted on glasses, (ii) with initial processing providing direction-of-arrival estimation of a desired speaker using a direct-path dominance test robust to reverberation, combined with speaker separation for improved dynamic localization, (iii) followed by speech enhancement against interfering speakers and noise, (iv) and ends with applying binaural signal matching for headphone listening. This paper compares model-based processing to learning-based processing in both noisy and dynamic scenarios, and presents a novel processing using data from a real wearable array, studied by simulation and a listening test.
Hanan Beit-On, Moti Lugasi, Lior Madmoni, Anjali Menon, Anurag Kumar 0003, Jacob Donley, Vladimir Tourbabin, Boaz Rafaely
ICASSP8
2022 Spatial Audio Signal Enhancement by a Two-Stage Source - System Estimation With Frequency Smoothing for Improved Perception
abstract
In many applications, such as hearing aids and virtual reality, spatial audio is used to provide a more natural experience to the users. However, when captured in the real world, the audio signals may suffer from noise and interference. In this case, the challenge is to attenuate the undesired signals, while preserving the desired signals with their spatial information. In this paper, an approach for spatial signal enhancement is presented. This approach is based on two phases of estimation. The first phase is source signal estimation using a beamformer. Then, in the second phase, the acoustic transfer function (ATF) between the source and the array is estimated; this leads to an enhanced estimation of the desired signal at the microphones. This approach has been previously proposed, but was not investigated in depth. In this paper, a model for the estimated desired signals is developed. In contrast to other methods of spatial enhancement, no trade-off between noise reduction and signal distortion is found in this model, when there is a single desired source and a single interfering source in a reverberant room. To overcome the limited accuracy of ATF estimation for short duration signals, frequency smoothing is applied. Listening tests verify the performance of the proposed approach.
Moti Lugasi, Anjali Menon, Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Applied Methods for Sparse Sampling of Head-Related Transfer Functions
abstract
Production of high fidelity spatial audio applications requires individual head-related transfer functions (HRTFs). As the acquisition of HRTF is an elaborate process, interest lies in interpolating full length HRTF from sparse samples. Ear-alignment is a recently developed pre-processing technique, shown to reduce an HRTF’s spherical harmonics order, thus permitting sparse sampling over fewer directions. This paper describes the application of two methods for ear-aligned HRTF interpolation by sparse sampling: Orthogonal Matching Pursuit and Principal Component Analysis. These methods consist of generating unique vector sets for HRTF representation. The methods were tested over an HRTF dataset, indicating that interpolation errors using small sampling schemes may be further reduced by up to 5 dB in comparison with spherical harmonics interpolation.
Lior Arbel, Zamir Ben-Hur, David L. Alon, Boaz Rafaely
ICASSP4
2021 Blind Amplitude Estimation of Early Room Reflections Using Alternating Least Squares
abstract
Estimation of the properties of early room reflections is an important task in audio signal processing, with applications in beamforming, source separation, room geometry inference, and spatial audio. While methods exist to blindly estimate the direction of arrival and delay of the early reflections, the blind estimation of reflection amplitudes remains an open problem. This work presents a preliminary attempt to blindly estimate reflection amplitudes. An iterative estimator is suggested, based on maximum likelihood and alternating least squares. We discuss some fundamental scaling ambiguities of the problem, and show connections between the proposed method and raking beamformers. A simulation study demonstrates the effectiveness of the proposed method.
Tom Shlomo, Boaz Rafaely
ICASSP2
2021 Binaural Reproduction Based on Bilateral Ambisonics and Ear-Aligned HRTFs
abstract
Reproduction of high quality spatial sound has gained considerable importance with the recent technology developments in the fields of virtual and augmented reality. Recently, the reproduction of binaural signals in the Spherical-Harmonics (SH) domain has been proposed. This is performed by using SH representations of the sound-field and the Head-Related Transfer Function (HRTF). These processes offer the flexibility to control the reproduced binaural signals, by manipulating the sound-field or the HRTFs using algorithms that operate directly in the SH domain. However, in most practical cases, the binaural reproduction is order-limited, which introduces truncation error that has a detrimental effect on the perception of the reproduced signals, mainly due to the truncation of the HRTF. A recent study showed that pre-processing of the HRTF by ear-alignment reduces its effective SH order, which may be beneficial for alleviating the above effect. In this paper, a method to incorporate the pre-processed ear-aligned HRTF into the binaural reproduction process is presented. The method uses Ambisonics representation of the sound-field formulated at the two ears, and is denoted here as Bilateral Ambisonics. The proposed method leads to a significant reduction in errors due to the limited-order reproduction, which yields a substantial improvement in perceived binaural reproduction quality even with SH as low as first order.
Zamir Ben-Hur, David L. Alon, Ravish Mehra, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 The Effect of Partial Time-Frequency Masking of the Direct Sound on the Perception of Reverberant Speech
abstract
The perception of sound in real-life acoustic environments, such as enclosed rooms or open spaces with reflective objects, is affected by reverberation. Hence, reverberation is extensively studied in the context of auditory perception, with many studies highlighting the importance of the direct sound for perception. Based on this insight, speech processing methods often use time-frequency (TF) analysis to detect TF bins that are dominated by the direct sound, and then use the detected bins to reproduce or enhance the speech signals. The detection of bins dominated by the direct sound is typically based on an objective measure, such as the direct-to-reverberant ratio (DRR). However, the relation between the DRR in the TF bins and the spatial perception of the reverberant sound which is reproduced from these bins is still not clear. It is the aim of this paper to provide some insights into this relation, specifically for reverberant speech, focusing on bins with high DRR. This is performed using a listening experiment, where high DRR bins within a reverberant speech signal have been masked in the TF domain, based on various DRR thresholds. The results show that the percentage of high-DRR TF bins that were masked may better indicate the quality of spatial perception, compared to the specific value of the DRR threshold. The insights from this work could be incorporated into spatial audio techniques that reproduce the direct sound of reverberant speech, and potentially improve spatial perception. This was illustrated with an implementation of directional audio coding that was studied with an additional listening experiment supporting the previously described results.
Lior Madmoni, Shir Tibor, Israel Nelken, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Robustness of Acoustic Rake Filters in Minimum Variance Beamforming
abstract
Acoustic rake filters perform coherent summation of the early room reflections using beamforming, with the aim of improving beamforming performance. This concept has been investigated for speech enhancement applications, improving noise reduction and late reverberation attenuation. Current studies typically assume that the parameters of the early reflections, such as the direction-of-arrival, delay and amplitude, are known in advance in the rake filter design. This work presents a novel investigation of the acoustic rake filter in a more practical context, focusing on the minimum variance distortionless response (MVDR) formulation. First, the sensitivity of the filter performance to perturbations in the reflection parameters is derived analytically, and investigated numerically using Monte Carlo simulations with a spherical microphone array. Then, an end-to-end example of rake filtering in a blind scenario is presented, where the reflection parameters are estimated from speech signals without any prior information. This example demonstrates for the first time the use of rake filtering in a realistic scenario.
Ran Weisman, Tom Shlomo, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 The Importance of Time-Frequency Averaging for Binaural Speaker Localization in Reverberant Environments
Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
INTERSPEECH3
2020 Spatial Covariance Matrix Estimation for Reverberant Speech with Application to Speech Enhancement
Ran Weisman, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely
INTERSPEECH4
2020 Focusing and Frequency Smoothing for Arbitrary Arrays With Application to Speaker Localization
abstract
The coherent signal subspace method (CSSM) enables the direction-of-arrival (DoA) estimation of coherent sources with subspace localization methods. The focusing process that aligns the signal subspaces within a frequency band to its central frequency is central to the CSSM. Within current focusing approaches, a direction-independent focusing approach may be more suitable for reverberant environments since no initial estimation of the sources' DoAs is required. However, these methods use integrals over the steering function, and cannot be directly applied to arrays around complex scattering structures, such as robot heads. In this article, current direction-independent focusing methods are extended to arrays for which the steering function is available only for selected directions, typically in a numerical form. Spherical harmonics decomposition of the steering function is then employed to formulate several aspects of the focusing error. A case of two coherent sources is studied and guidelines for the selection of the frequency smoothing bandwidth are suggested. The performance of the proposed methods is then investigated for an array that is mounted on a robot head. The focusing process is integrated within the direct-path dominance (DPD) test method for speaker localization, originally designed for spherical arrays, extending its application to arrays with arbitrary configurations. Finally, experiments with real data verify the feasibility of the proposed method to successfully estimate the DoAs of multiple speakers under real-world conditions.
Hanan Beit-On, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Speech Enhancement Using Masking for Binaural Reproduction of Ambisonics Signals
abstract
Speech enhancement in a single channel has been well studied in the literature in applications such as speech communication systems. However, in emerging applications such as virtual reality and spatial audio, in addition to attenuating undesired signals, the ability to preserve the spatial information of the desired signal captured in a noisy environment is of great importance. Nevertheless, there are only a few studies in the literature that propose solutions to this challenge. Most of these studies present solutions that attenuate the undesired signals, while preserving only limited spatial information regarding the desired signal, such as the direction of arrival (DOA). Methods that preserve complete spatial information have only recently been suggested, and have not been studied comprehensively. In this paper, two such methods based on time-frequency masking are investigated with the aim of attenuating the undesired signal, while preserving the spatial components of the desired signal. The first is referred to as spatial masking and is based on masking in the plane wave density (PWD) domain, and the second on masking in the spherical harmonics (SH) domain. The two methods are compared with a reference method, based on beamforming followed by single-channel time-frequency masking. Objective analysis and two listening tests were conducted in order to evaluate the performance of these methods for speech enhancement. It was shown that the spatial masking based method better preserves the desired component of the sound field, while the performance of the SH based method more strongly depends on the sources' distances. On the other hand, the SH based method better preserves the DOA of the residual noise, while the DOA of the residual noise under the spatial masking based method is strongly affected by the undesired signal.
Moti Lugasi, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Perceptually-Transparent Online Estimation of Two-Channel Room Transfer Function for Sound Calibration
abstract
Sound calibration is employed in many commercial audio systems for improving sound quality. This process includes the estimation of the room transfer function (RTF) between each loudspeaker and a microphone located at the listeners' position. Current methods for RTF estimation employ calibration signals, such as noise or tones, in a dedicated process applied to each loudspeaker separately. Such an estimation disrupts normal playback, is time consuming, and requires user intervention. A perceptually-transparent online RTF estimation method for a two-channel system, which employs calibration signals generated using the original audio signals and complementary filters, is proposed in this article. These calibration signals are uncorrelated across the two channels, which facilitates the online estimation of both channels using a single microphone. Estimation performance is investigated for an experimental system and displays low estimation errors. Finally, a subjective evaluation via a listening test shows that playback of calibration signals is perceptually-transparent under some of the conditions investigated.
Hai Morgenstern, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Efficient Representation and Sparse Sampling of Head-Related Transfer Functions Using Phase-Correction Based on Ear Alignment
abstract
With the proliferation of high quality virtual reality systems, the demand for high fidelity spatial audio reproduction has grown. This requires individual head-related transfer functions (HRTFs) with high spatial resolution. Acquiring such HRTFs is not always possible, which motivates the need for sparsely sampled HRTFs. Additionally, real-time applications require compact representation of HRTFs. Recently, spherical-harmonics (SH) has been suggested for efficient interpolation and representation of HRTFs. However, representation of sparse HRTFs with a limited SH order may introduce spatial aliasing and truncation errors, which have a detrimental effect on the reproduced spatial audio. This is because the HRTF is inherently of a high spatial order. One approach to overcome this limitation is to pre-process the HRTF, with the aim of reducing its effective SH order. A recent study showed that order-reduction can be achieved by time-alignment of HRTFs, through numerical estimation of the time delays of the HRTFs. In this paper, a new method for pre-processing HRTFs in order to reduce their effective order is presented. The method uses phase-correction based on ear alignment, by exploiting the dual-centering nature of HRTF measurements. In contrast to time-alignment, the phase-correction is performed parametrically, making it more robust to measurement noise. The SH order reduction and ensuing interpolation errors due to sparse sampling were analyzed for these two methods. Results indicate significant reduction in the effective SH order, where only 100 measurements and order 6 are required to achieve a normalized mean square error below -10 dB compared to a fully-sampled, high-order HRTF.
Zamir Ben-Hur, David L. Alon, Ravish Mehra, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Sparse Head-Related Transfer Function Representation with Spatial Aliasing Cancellation
abstract
High-fidelity 3D audio experience requires accurate individual head-related transfer function (HRTF) representation. However, the process of measuring individual HRTFs typically involves measurements from hundreds of directions, with specialized and expensive equipment, which makes this process inaccessible for most users. In this paper, a new technique to reconstruct high resolution individual HRTFs from sparse measurements is presented. This is achieved by minimizing the spatial aliasing error in the spherical harmonics (SH) representation of the HRTFs, and by incorporating statistics calculated from a set of reference HRTFs, leading to an optimal minimum mean-square error solution. A quantitative analysis of the proposed method illustrates its benefits even for extreme cases, such as using only 25 individual HRTF measurements and a generic HRTF as a reference.
David L. Alon, Zamir Ben-Hur, Boaz Rafaely, Ravish Mehra
ICASSP3
2018 Speaker localization using direct path dominance test based on sound field directivity
Boaz Rafaely, Koby Alhaiany
Signal Process.1
2017 Speaker localization in reverberant rooms based on direct path dominance test statistics
abstract
Speaker localization using microphone arrays is typically based on the expected phase and amplitude differences between microphones as a function of the wave arrival direction. However, in rooms with significant reverberation, the direct sound is contaminated by reflections and localization often fails. Recently, a reverberation-robust localization method was proposed, which uses only the direct-path bins in the short-time Fourier transform (STFT) of the speech signals. The method is based on thresholding according to the ratio between the first two singular values of the spatial spectrum matrix. In this work, a confidence measure is developed based on this ratio, which is then used for speaker localization in a statistical estimation framework, based on a Gaussian mixture model. The paper presents the theory of the proposed method and simulation examples validating the advantages of the new approach.
Boaz Rafaely, Dorothea Kolossa
ICASSP1
2017 Efficient inverse spatially localized spherical Fourier transform with kernel partitioning of the sphere
Uri Abend, Boaz Rafaely
Signal Process.2
2016 Beamforming with Optimal Aliasing Cancellation in Spherical Microphone Arrays
abstract
Spherical microphone arrays facilitate three-dimensional processing and analysis of sound fields in applications such as music recording, beamforming and room acoustics. The frequency bandwidth of operation is constrained by the array configuration. At high frequencies, spatial aliasing leads to side-lobes in the array beam pattern, which limits array performance. Previous studies proposed increasing the number of microphones or changing other characteristics of the array configuration to reduce the effect of aliasing. In this paper we present a method to design beamformers that overcome the effect of spatial aliasing by suppressing the undesired side-lobes through signal processing without physically modifying the configuration of the array. This is achieved by modeling the expected aliasing pattern in a maximum-directivity beamformer design, leading to a higher directivity index at frequencies previously considered to be out of the operating bandwidth, thereby extending the microphone array frequency range of operation. Aliasing cancellation is then extended to other beamformers. A simulation example with a 32-element spherical microphone array illustrates the performance of the proposed method. An experimental example validates the theoretical results in practice.
David L. Alon, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Theory and Perceptual Evaluation of the Binaural Reproduction and Beamforming Tradeoff in the Generalized Spherical Array Beamformer
abstract
Microphone arrays are widely used in speech enhancement systems for noisy and reverberant environments. Recently, a generalized spherical array beamforming approach was developed incorporating binaural sound reproduction in the beamforming process. This generalized spherical array beamformer (GSB) maintains the spatial information through the binaural cues and improves both the spatial realism and the speech intelligibility. In this paper, the theory of the tradeoff that arises when incorporating both beamforming and binaural reproduction in a single array is developed and investigated through a simulation study and a listening test. By representing the GSB formulation in matrix form for investigating the single plane-wave scenario, two measures are developed in order to evaluate the performance of the GSB in terms of both binaural reproduction and spatial selectivity. These measures are then employed in the evaluation of the performance of various GSB beam-patterns using simulations. A listening test experiment that validates the simulation results is then reported. Results validate the theory, i.e., the GSB can be used to integrate successfully binaural reproduction and beamforming, allowing the user to emphasize either of the two, but with a clear tradeoff; improving one is only possible at the expense of degrading the other.
Michael Jeffet, Noam R. Shabtai, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Enhanced robot audition by dynamic acoustic sensing in moving humanoids
abstract
Auditory systems of humanoid robots usually acquire the surrounding sound field by means of microphone arrays. These arrays can undergo motion related to the robot's activity. The conventional approach to dealing with this motion is to stop the robot during sound acquisition. This approach avoids changing the positions of the microphones during the acquisition and reduces the robot's ego-noise. However, stopping the robot can interfere with the naturalness of its behaviour. Moreover, the potential performance improvement due to motion of the sound acquiring system can not be attained. This potential is analysed in the current paper. The analysis considers two different types of motion: (i) rotation of the robot's head and (ii) limb gestures. The study presented here combines both theoretical and numerical simulation approaches. The results show that rotation of the head improves the high-frequency performance of the microphone array positioned on the head of the robot. This is complemented by the limb gestures, which improve the low-frequency performance of the array positioned on the torso and limbs of the robot.
Vladimir Tourbabin, Hendrik Barfuss, Boaz Rafaely, Walter Kellermann
ICASSP3
2015 Binaural Reproduction of Finite Difference Simulations Using Spherical Array Processing
abstract
Due to its efficiency and simplicity, the finite-difference time-domain method is becoming a popular choice for solving wideband, transient problems in various fields of acoustics. So far, the issue of extracting a binaural response from finite difference simulations has only been discussed in the context of embedding a listener geometry in the grid. In this paper, we propose and study a method for binaural response rendering based on a spatial decomposition of the sound field. The finite difference grid is locally sampled using a volumetric array of receivers, from which a plane wave density function is computed and integrated with free-field head related transfer functions, in the spherical harmonics domain. The volumetric array is studied in terms of numerical robustness and spatial aliasing. Analytic formulas that predict the performance of the array are developed, facilitating spatial resolution analysis and numerical binaural response analysis for a number of finite difference schemes. Particular emphasis is placed on the effects of numerical dispersion on array processing and on the resulting binaural responses. Our method is compared to a binaural simulation based on the image method. Results indicate good spatial and temporal agreement between the two methods.
Jonathan Sheaffer, Maarten van Walstijn, Boaz Rafaely, Konrad Kowalczyk
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots
abstract
The auditory system of humanoid robots has gained increased attention in recent years. This system typically acquires the surrounding sound field by means of a microphone array. Signals acquired by the array are then processed using various methods. One of the widely applied methods is direction of arrival estimation. The conventional direction of arrival estimation methods assume that the array is fixed at a given position during the estimation. However, this is not necessarily true for an array installed on a moving humanoid robot. The array motion, if not accounted for appropriately, can introduce a significant error in the estimated direction of arrival. The current paper presents a signal model that takes the motion into account. Based on this model, two processing methods are proposed. The first one compensates for the motion of the robot. The second method is applicable to periodic signals and utilizes the motion in order to enhance the performance to a level beyond that of a stationary array. Numerical simulations and an experimental study are provided, demonstrating that the motion compensation method almost eliminates the motion-related error. It is also demonstrated that by using the motion-based enhancement method it is possible to improve the direction of arrival estimation performance, as compared to that obtained when using a stationary array.
Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Localization of Multiple Speakers under High Reverberation using a Spherical Microphone Array and the Direct-Path Dominance Test
abstract
One of the major challenges encountered when localizing multiple speakers in real world environments is the need to overcome the effect of multipath distortion due to room reverberation. A wide range of methods has been proposed for speaker localization, many based on microphone array processing. Some of these methods are designed for the localization of coherent sources, typical of multipath environments, and some have even reported limited robustness to reverberation. Nevertheless, speaker localization under conditions of high reverberation still remains a challenging task. This paper proposes a novel multiple-speaker localization technique suitable for environments with high reverberation, based on a spherical microphone array and processing in the spherical harmonics (SH) domain. The non-stationarity and sparsity of speech, as well as frequency smoothing in the SH domain, are exploited in the development of a direct-path dominance test. This test can identify time-frequency (TF) bins that contain contributions from only one significant source and no significant contribution from room reflections, such that localization based on these selected TF-bins is performed accurately, avoiding the potential distortion due to other sources and reverberation. Computer simulations and an experiment in a real reverberant room validate the robustness of the proposed method in the presence of high reverberation .
O. Nadiri, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Generalized Spherical Array Beamforming for Binaural Speech Reproduction
abstract
Microphone arrays are used in speech signal processing applications such as teleconferencing and telepresence, in order to enhance a desired speech signal in the presence of speech signals from other speakers, reverberation and background noise. These arrays usually provide a single-channel output, so that no spatial information is available in the output signal. However, spatial information on the sound sources may increase the intelligibility of a speech signal perceived by a human listener. This work presents a mathematical framework for generalized spherical array beamforming that in addition to suppressing noise and reverberation, is aiming to preserve spatial information on the sources in the recording venue. The generalized beamforming, formulated in the spherical harmonics domain, is based on binaural sound reproduction where the head-related transfer functions are incorporated into a headphones presentation. The performance of the proposed generalized beamformer is compared to that of a single-channel output maximum-directivity beamformer. Listening tests with human subjects show that when the generalized beamformer is used the intelligibility is improved at low input SNRs.
Noam R. Shabtai, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Theoretical framework for the optimization of microphone array configuration for humanoid robot audition
abstract
An important aspect of a humanoid robot is audition. Previous work has presented robot systems capable of sound localization and source segregation based on microphone arrays with various configurations. However, no theoretical framework for the design of these arrays has been presented. In the current paper, a design framework is proposed based on a novel array quality measure. The measure is based on the effective rank of a matrix composed of the generalized head related transfer functions (GHRTFs) that account for microphone positions other than the ears. The measure is shown to be theoretically related to standard array performance measures such as beamforming robustness and DOA estimation accuracy. Then, the measure is applied to produce sample designs of microphone arrays. Their performance is investigated numerically, verifying the advantages of array design based on the proposed theoretical framework.
Vladimir Tourbabin, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Spherical loudspeaker array beamforming in enclosed sound fields by MIMO optimization
abstract
Spherical loudspeaker arrays have been recently studied for a variety of applications such as spatial sound reproduction and room acoustics. Array directivity is typically designed for free field, although loudspeaker arrays often operate in enclosures. Therefore, current methods for directivity design may not be suitable. We present a new method for designing loudspeaker array directivity in enclosures, where the direct sound measured by a microphone array is emphasized compared to room reflections. The two spherical arrays form an acoustic multiple-input multiple-output (MIMO) system for which the loudspeaker array beamforming coefficients are designed based on the system transfer matrix. The proposed method uses the average output of the microphone array, over all look directions to avoid signal cancellation due to room reflections. The performance of the proposed method is studied through a simulation example.
Hai Morgenstern, Boaz Rafaely
ICASSP2
2013 Binaural sound reproduction beamforming using spherical microphone arrays
abstract
Currently employed microphone arrays usually have a single-channel output, such that no spatial information can be perceived by a human listener. However, spatial information may trigger spatial mechanisms in the human auditory system which can improve the intelligibility. This work presents a mathematical framework for the binaural beamforming approach for the ideal and order-limited representation in the spherical harmonics domain. The performance of the proposed binaural beamformer is compared to that of a monaural maximum directivity beamformer using objective signal-based measures and subjective listening tests. It is shown that using the binaural beamformer results in higher intelligibility than the monaural beamformer.
Noam R. Shabtai, Boaz Rafaely
ICASSP2
2013 Theoretical framework for the design of microphone arrays for robot audition
abstract
An important part of a human-like robot is robot audition. Previous work presented systems capable of sound localization and source segregation based on microphone arrays of various configurations. However, no theoretical framework for assessing the quality of these array configurations has been presented. In the current paper such a measure is proposed based on the generalized HRTFs that account for microphone positions other than the ears. The measure is analyzed theoretically with respect to beamforming robustness and DOA estimation accuracy. The measure is then used to find the optimal location of a single microphone and a pair of microphones based on the generalized HRTF database obtained by means of BEM simulation. The results are not surprising, showing that the best position of a single microphone is the ear canal. For a pair of microphones, the results generally show that the sensors should be maximally spatially separated.
Vladimir Tourbabin, Boaz Rafaely
ICASSP2
2013 Linearly-Constrained Minimum-Variance Method for Spherical Microphone Arrays Based on Plane-Wave Decomposition of the Sound Field
abstract
Speech signals recorded in real environments may be corrupted by ambient noise and reverberation. Therefore, noise reduction and dereverberation algorithms for speech enhancement are typically employed in speech communication systems. Although microphone arrays are useful in reducing the effect of noise and reverberation, existing methods have limited success in significantly removing both reverberation and noise in real environments. This paper presents a method for noise reduction and dereverberation that overcomes some of the limitations of previous methods. The method uses a spherical microphone array to achieve plane-wave decomposition (PWD) of the sound field, based on direction-of-arrival (DOA) estimation of the desired signal and its reflections. A multi-channel linearly-constrained minimum-variance (LCMV) filter is introduced to achieve further noise reduction. The PWD beamformer achieves dereverberation while the LCMV filter reduces the uncorrelated noise with a controllable dereverberation constraint. In contrast to other methods, the proposed method employs DOA estimation, rather than room impulse response identification, to achieve dereverberation, and relative transfer function (RTF) estimation between the source reflections to achieve noise reduction while avoiding signal cancellation. The paper includes a simulation investigation and an experimental study, comparing the proposed method to currently available methods.
Yotam Peled, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.2
2012 Analysis of acoustic MIMO systems in enclosed sound fields
abstract
Methods for room acoustic analysis based on a single loudspeaker and a single microphone have been extensively studied, both theoretically and experimentally. Although measurement techniques for acquiring and processing spatial information in a room using arrays of loudspeakers and microphones have been recently proposed, the current literature in room acoustics does not describe the use of multiple-input multiple-output (MIMO) systems in a comprehensive manner. The aim of this paper is, therefore, to present an initial theoretical framework for the spatial analysis of enclosed sound fields using an acoustic MIMO system, based on a spherical loudspeaker array and a spherical microphone array. The fundamental characteristics of the system transfer matrix are shown to be invariant to rotation of the arrays, with its rank dependant on the number of reflections in the room. The paper concludes with a simulation study validating the theoretical models.
Hai Morgenstern, Boaz Rafaely
ICASSP2
2012 Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays
abstract
One of the uses of sensor arrays is for spatial filtering or beamforming. Current digital signal processing methods facilitate complex-weighted beamforming, providing flexibility in array design. Previous studies proposed the use of real-valued beamforming weights, which although reduce flexibility in design, may provide a range of benefits, e.g., simplified beamformer implementation or efficient beamforming algorithms. This paper presents a new method for the design of arrays with real-valued weights, that achieve maximum directivity, providing closed-form solution to array weights. The method is studied for linear and spherical arrays, where it is shown that rigid spherical arrays are particularly suitable for real-weight designs as they do not suffer from grating lobes, a dominant feature in linear arrays with real weights. A simulation study is presented for linear and spherical arrays, along with an experimental investigation, validating the theoretical developments.
Vladimir Tourbabin, Morag Agmon, Boaz Rafaely, Joseph Tabrikian
IEEE Trans. Speech Audio Process.3
2011 Near-Field Spherical Microphone Array Processing With Radial Filtering
abstract
This paper presents an analysis of spherical microphone array capabilities in the near-field, with an emphasis on radial filtering of sources in a given direction. The near-field of the array is defined in terms of frequency and distance from the array. Directional beamforming is demonstrated given the near-field radial compensation filter, which yields a desired directional beampattern at a chosen distance from the array. This pattern deteriorates as the source draws away from the array. Next, a framework is presented for radial filter design, enabling distance discrimination between sources positioned in the same direction relative to the array. Design examples include Dolph-Chebyshev radial filtering, radial notch filtering, and numerical design. Performance is analyzed in terms of spatial response and robustness to noise. Results show radial filtering is practical for improving attenuation of far-field and near-field interfering sources relative to a desired source positioned in the same direction.
Etan Fisher, Boaz Rafaely
IEEE Trans. Speech Audio Process.2
2011 Bessel Nulls Recovery in Spherical Microphone Arrays for Time-Limited Signals
abstract
Spherical microphone arrays have recently been developed for a wide range of applications. In particular, scanning arrays in open-sphere configuration have been employed for room acoustics analysis, based on room impulse response measurements. However, it has been shown that the simple single-sphere configuration suffers from ill-conditioning around the zeros of the spherical Bessel function, and so alternative configurations, typically more complex, such as dual-sphere and single-sphere with cardioid microphones, have been proposed to overcome the effect of the Bessel nulls. This paper shows that in the particular case of time-limited signals, for example in room impulse response measurements, the ill-conditioning due to the spherical Bessel functions can be reduced using nonuniform sampling in the frequency domain. Following a presentation of the theory and a simulation study, an experimental example is presented for use of the method in a single-sphere measurement in an auditorium.
Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.1
2011 Optimal Model-Based Beamforming and Independent Steering for Spherical Loudspeaker Arrays
abstract
Spherical loudspeaker arrays have been recently studied for directional sound radiation, where the compact arrangement of the loudspeaker units around a sphere facilitated the control of sound radiation in three-dimensional space. Directivity of sound radiation, or beamforming, was achieved by driving each loudspeaker unit independently, where the design of beamforming weights was typically achieved by numerical optimization with reference to a given desired beam pattern. This is in contrast to the methods already developed for microphone arrays in general and spherical microphone arrays in particular, where beamformer weights are designed to satisfy a wider range of objectives, related to directivity, robustness, and side-lobe level, for example. This paper presents the development of a physical-model-based, optimal beamforming framework for spherical loudspeaker arrays, similar to the framework already developed for spherical microphone arrays, facilitating efficient beamforming in the spherical harmonics domain, with independent steering. In particular, it is shown that from a beamforming perspective, the spherical loudspeaker array is similar to the spherical microphone array with microphones arranged around a rigid sphere. Experimental investigation validates the theoretical framework of beamformer design.
Boaz Rafaely, Dima Khaykin
IEEE Trans. Speech Audio Process.1
2010 Method for dereverberation and noise reduction using spherical microphone arrays
abstract
A method for dereverberation and noise reduction is presented. The method is designed for a spherical microphone array, and formulated in the spherical harmonics domain, based on an acoustic model that is also formulated in the spherical harmonics domain. The novelty in the proposed method is the dereverberation process, which exploits the useful formulation in the spherical harmonics domain, which facilitates dereverberation by employing DOA estimation, rather that room impulse response identification. Noise reduction is further performed by a linearly constrained minimum variance filter, where the array output power is minimized with constraint of distortionless response to the direct sound. The paper concludes with a simulation investigation and comparison to theoretical results.
Yotam Peled, Boaz Rafaely
ICASSP2
2010 Golden-Ratio Sampling for Scanning Circular Microphone Arrays
abstract
Circular microphone arrays have been recently studied for music recordings, beamforming, and sound-field analysis. Scanning circular arrays that use a single microphone to acquire measurements at a range of spatial sampling positions have been employed to gather an entire array data. Uniform sampling along the circle is typically used, but this scheme may not allow experimental flexibility. If, for any reason, the experiment was not completed in full, the data gathered may not be sufficient to perform the required processing. This paper presents an alternative spatial sampling scheme for scanning circular microphone arrays, which is more robust against uncertainty in the duration of the experiment. The proposed spatial sampling method is based on fixed-angle steps along the circle defined by the golden-ratio. Theoretical analysis of the golden-ratio sampling method is presented, supported by simulation examples, showing that experimental flexibility is achieved, but at the expanse of some reduction in numerical robustness.
Maor Kleider, Boaz Rafaely, Barak Weiss, Eitan Bachmat
IEEE Trans. Speech Audio Process.2
2008 The nearfield spherical microphone array
abstract
A nearfield spherical microphone array is presented. The nearfield criterion of the spherical array is defined in terms of array order, frequency and location. It is shown that given a source in the nearfield, significant attenuation of farfield interference is achieved. Also, nearfield sources may be attenuated relative to farfield sources. Dereverberation of a nearfield source in a reverberant enclosure is demonstrated using the nearfield microphone array.
Etan Fisher, Boaz Rafaely
ICASSP2
2008 Reverberation matching for speaker recognition
abstract
Speech recorded by a distant microphone in a room may be subject to reverberation. Performance of a speaker verification system may degrade significantly for reverberant speech, with severe consequences in a wide range of real applications. This paper presents a comprehensive study of the effect of reverberation on speaker verification, and investigates approaches to reduce the effect of reverberation: training target models with reverberant speech signals and using acoustically matched models for the reverberant speech under test, score normalization methods to improve the reverberation robustness, and also reverberation classification via the background model scores. Experimental investigation is performed, using simulated and measured room impulse responses, NIST-based speech database, and AGMM based speaker verification system, showing significant improvement in performance.
Itai Peer, Boaz Rafaely, Yaniv Zigel
ICASSP2
2008 Spherical microphone array with multiple nulls for analysis of directional room impulse responses
abstract
A spherical microphone array is presented which incorporates multiple nulls in the beampattern for analysis of directional room impulse responses. Improved performance is achieved compared to spherical arrays with regular beampatterns due to the ability to attenuate undesired room reflections. Formulation of the multiple-null spherical array processing is presented, both in the space and spherical harmonics domains. The paper concludes with experimental investigation using impulse response data measured in an auditorium.
Boaz Rafaely
ICASSP1
2008 Spherical Microphone Array Beam Steering Using Wigner-D Weighting
abstract
Spherical microphone arrays, which have been recently developed and proposed for various applications, typically employ beam patterns that are rotationally symmetric about the look direction, providing efficient beam steering in the spherical harmonics domain. However, in some situations, a more general beam pattern may be desired. This letter presents the theory and a simulation example for steering general beam patterns in spherical microphone arrays. Beam steering, formulated as a rotation of the beam pattern, is achieved by weighting the beam pattern coefficients in the spherical harmonics domain with Wigner-D functions that hold the rotation angles as parameters. A matrix formulation is provided, with successive rotations formulated as matrix products.
Boaz Rafaely, Maor Kleider
IEEE Signal Process. Lett.1
2008 The Spherical-Shell Microphone Array
abstract
Spherical microphone arrays have been recently studied for a wide range of applications. In particular, microphones arranged around an open or virtual sphere are useful in scanning microphone arrays for sound field analysis. However, open-sphere spherical arrays have been shown to have poor robustness at frequencies related to the zeros of the spherical Bessel functions. This paper presents a framework for the analysis of array robustness using the condition number of a given matrix, and then proposes several robust array configurations. In particular, a dual-sphere configuration previously presented which uses twice as many microphones compared to a single-sphere configuration is analyzed. This paper then shows that high robustness can be achieved without increasing the number of microphones by arranging the microphones in the volume of a spherical shell. Another simpler configuration employs a single sphere and an additional microphone at the sphere center, showing improved robustness at the low-frequency range. Finally, the white-noise gain of the arrays is investigated verifying that improved white-noise gain is associated with lower matrix condition number.
Boaz Rafaely
IEEE Trans. Speech Audio Process.1
2007 Open-Sphere Designs for Spherical Microphone Arrays
abstract
Spherical microphone arrays have been studied for a wide range of applications, one of which is acoustic measurement and analysis. Since a minimal interaction between the array and the measured sound field is an advantage in this case, open-sphere arrays are preferable compared to rigid-sphere arrays. However, it has been shown that open-sphere arrays suffer from numerical ill-conditioning at frequencies which correspond to the nodal values of the spatial spherical modes with the result of excessive noise at these frequencies. A method for overcoming this problem using an open dual-sphere array is proposed in this correspondence and then investigated and compared to an array configured around a rigid sphere and an array composed of cardioid microphones. An optimal value for the ratio of the two spheres is derived, and simulation examples illustrating the advantage of the dual-sphere array are finally presented
I. Balmages, Boaz Rafaely
IEEE Trans. Speech Audio Process.2
2006 Unbiased adaptive feedback cancellation in hearing aids by closed-loop identification
abstract
Hearing aid users often suffer from feedback which causes "howling" and limits the maximum stable gain. The direct method of adaptive feedback cancellation is widely used to mitigate feedback, however it is less effective for high forward path gains due to the inherent bias in the feedback path estimate. This bias can be reduced at the expense of artificial delays which can potentially introduce pre-echo and "comb filter" effects. The direct method also tends to cancel tonal audio signals such as alarms and music. We propose using a system identification method in closed-loop for unbiased feedback cancellation which does not possess the negative characteristics manifested by the direct method, but uses an identification signal. The two-stage method which employs two adaptive filters, one to identify the entire closed-loop, and another to extract the feedback path response was the preferred unbiased method due to its low computation and good feedback identification particularly during "howling." This paper presents the theory of the proposed new method, and computer simulations which demonstrate its advantages over the widely used direct method.
Ngwa A. Shusina, Boaz Rafaely
IEEE Trans. Speech Audio Process.2
2005 Phase-mode versus delay-and-sum spherical microphone array processing
abstract
Phase-mode spherical microphone array processing, also known as spherical harmonic array processing, has been recently studied for various applications. The spherical array configuration provides desired three-dimensional symmetry, while the phase modes provide frequency-independent spatial processing. This letter employs the spherical harmonic framework to compare the well-known delay-and-sum to the phase-mode processing for spherical arrays. The two approaches show similar performance at frequencies where the upper spherical harmonic order equals the product of the wave number and sphere radius. However, at lower frequencies, phase-mode processing maintains the same directivity, limited by signal-to-noise ratio, while for delay-and-sum, spatial resolution deteriorates.
Boaz Rafaely
IEEE Signal Process. Lett.1
2005 Analysis and design of spherical microphone arrays
abstract
Spherical microphone arrays have been recently studied for sound-field recordings, beamforming, and sound-field analysis which use spherical harmonics in the design. Although the microphone arrays and the associated algorithms were presented, no comprehensive theoretical analysis of performance was provided. This work presents a spherical-harmonics-based design and analysis framework for spherical microphone arrays. In particular, alternative spatial sampling schemes for the positioning of microphones on a sphere are presented, and the errors introduced by finite number of microphones, spatial aliasing, inaccuracies in microphone positioning, and measurement noise are investigated both theoretically and by using simulations. The analysis framework can also provide a useful guide for the design and analysis of more general spherical microphone arrays which do not use spherical harmonics explicitly.
Boaz Rafaely
IEEE Trans. Speech Audio Process.1
2003 Robust compensation with adaptive feedback cancellation in hearing aids
Boaz Rafaely, Ngwa A. Shusina, Joanna L. Hayes
Speech Commun.1
2000 Control of feedback in hearing aids-a robust filter design approach
abstract
A bound on the variability of the feedback path is employed in the design of fixed FIR hearing aid filters that are robust to the specified variability, thus avoiding instability and howling in everyday use. A design example is presented for a linear gain hearing aid filter with a given maximal mismatch of the feedback cancellation filter.
Boaz Rafaely, Mariano Roccasalva-Firenze
IEEE Trans. Speech Audio Process.1
1997 Rapid frequency-domain adaptation of causal FIR filters
abstract
Normalizing the convergence coefficient of the block frequency-domain least mean square (LMS) algorithm in each frequency bin can improve the convergence rate, but in some applications can lead to a biased steady-state solution if the filter is constrained to be strictly causal. An algorithm is presented in which the spectral factors of the bin-normalized convergence coefficient are used before and after the causality constraint is applied in the adaptation algorithm, which converges rapidly to the optimal causal filter.
Stephen J. Elliott, Boaz Rafaely
IEEE Signal Process. Lett.2
1996 Audiometric ear canal probe with active ambient noise control
abstract
An audiometric earphone system is introduced, which employs both passive and active acoustic noise attenuation. The audiometric earphone includes a miniature speaker, an error microphone, and a reference microphone placed in a foam plug. It is inserted into a subject's ear canal during a hearing test, and reduces the ambient noise to bone conduction levels. Passive attenuation is achieved by the foam plug, and active attenuation is achieved by a feedforward active noise control implementation. Initially, the earphone was tested in mechanical models, and its causality and optimal characteristics were investigated. A causal controller was obtained only in a mechanical model with enlarged spacing between the reference microphone and the speaker. Up to 30 dB attenuation of broadband noise in the enlarged mechanical model and 40 dB attenuation of tones in the regular-size mechanical ear model were achieved. The system was then tested for its ability to reduce noise audibility by placing the audiometric earphone in the ear canal of human subjects. An attenuation of more than 30 dB of 500 Hz tone was achieved using a hearing threshold test. Although the system can currently attenuate only periodic noise when placed in the ear canal, with advanced technology it will be possible to design practical audiometric earphones for hearing tests.
Boaz Rafaely, Miriam Furst
IEEE Trans. Speech Audio Process.1