Wen Zhang 0002

dblp:43/2368-2 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-0752-6123ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Radiation and Directivity Analysis of a Vibrating Dome-Shaped Radiator Mounted on an Infinite Baffle
abstract
Accurate modeling and analysis of a radiator mounted on an infinite baffle are crucial for understanding its acoustic radiation characteristics. This paper investigates the radiation behavior of a convex dome-shaped radiator in such a condition, showing that, under the far-field approximation, the pressure field is the three-dimensional Fourier transform of the axisymmetric surface velocity distribution. The study includes a comparison of various typical velocity distributions in terms of their directivity factor (DF) and radiated sound power. Current velocity distributions often suffer from nulls in the DF, so we provide a detailed analysis to uncover the causes of this issue. To address this, we propose an equalization filter designed to smooth the DF across the entire frequency range. Simulations are performed to validate the theoretical findings and to showcase the improved performance of the proposed approach.
Junqing Zhang, Wen Zhang 0002, Jingdong Chen, Jacob Benesty
ICASSP2
2025 Conjugate Gradient and Variance Reduction Based Online ADMM for Low-Rank Distributed Networks
abstract
Modeling the relationships that may connect optimal parameter vectors is essential for the performance of parameter estimation methods in distributed networks. In this paper, we consider a low-rank relationship and introduce matrix factorization to promote this low-rank property. To devise a distributed algorithm that does not require any prior knowledge about the low-rank space, we first formulate local optimization problems at each node, which are subsequently addressed using the Alternating Direction Method of Multipliers (ADMM). Three subproblems naturally arise from ADMM, each resolved in an online manner with low computational costs. Specifically, the first one is solved using stochastic gradient descent (SGD), while the other two are handled using the conjugate gradient descent method to avoid matrix inversion operations. To further enhance performance, a variance reduction algorithm is incorporated into the SGD. Simulation results validate the effectiveness of the proposed algorithm.
Danqi Jin, Jie Chen 0022, Cédric Richard, Wen Zhang 0002
IEEE Signal Process. Lett.5
2025 Zeroth-Order Distributed Stochastic Optimization Over Riemannian Manifolds
abstract
Due to its ability to handle strict constraints on feasible domains, distributed optimization over a Riemannian manifold offers an attractive solution for many practical applications. To develop such an algorithm for scenarios where the explicit expression of the cost function is unavailable, we introduce the zeroth-order (ZO) Riemannian stochastic gradient into distributed optimization on a Riemannian manifold. Specifically, an intermediate estimate is first obtained through a local update step using the ZO Riemannian stochastic gradient, which is ap proximated based on two function evaluations. Subsequently, an improved estimate is derived by minimizing the weighted Fr´echet mean over the manifold using information from neighboring nodes. To further enhance performance, a mini-batch strategy is incorporated into the gradient estimation process. Finally, simulation results are presented to validate the effectiveness of the proposed algorithm.
Danqi Jin, Jie Chen 0022, Wen Zhang 0002
IEEE Signal Process. Lett.4
2024 Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split Network
abstract
Stereophonic music source separation (MSS) is a problem of extracting individual source tracks, e.g. bass, drums, vocals, from a stereo music recording. Deep neural network (DNN) based MSS systems have demonstrated great promise though spatial panning cues and time-frequency spectral structures in stereo music have not yet been fully explored in such systems and methods. This paper presents a spatially-informed MSS method using a bridging band-split neural network that incorporates both spatial and spectral information. The spatial panning angles of each target source are used as input of the network, along with the time-frequency spectrograms. Moreover, the inter-track correlations are exploited for further performance improvement. Experiments show that the proposed method outperforms significantly the baseline systems as the result of using spatial cues, spectral characteristics, and inter-track relationships.
Yichen Yang 0010, Xianrui Wang, Wen Zhang 0002, Shoji Makino, Jingdong Chen
ICASSP4
2024 Multi-Source DOA Estimation Using Higher-Order Pseudo Intensity Vector on a Spherical Microphone Array
abstract
Direction-of-arrival (DOA) estimation in environments with multiple sources and strong reverberation remains a great challenge. In this letter, we present a novel feature, the higher-order pseudo-intensity vector (HOPIV), derived from recordings obtained with a spherical microphone array. By exploiting the unique properties of the reactive intensity vector, which is derived from the HOPIV, we present a method to identify time-frequency points that are dominated by the direct path. We then propose a DOA estimation method that leverages the HOPIV's high spatial resolution for improving DOA estimation performance. Simulations and experiments show that the proposed method is able to yield superior performance compared to state-of-the-art techniques, even in highly reverberant environments.
Wen Zhang 0002, Jingdong Chen, Mengyao Zhu 0003, Chunjian Li
IEEE Signal Process. Lett.2
2024 Interference-Controlled Maximum Noise Reduction Beamformer Based on Deep-Learned Interference Manifold
abstract
Beamforming has been used in a wide range of applications to extract the signal of interest from microphone array observations, which consist of not only the signal of interest, but also noise, interference, and reverberation. The recently proposed interference-controlled maximum noise reduction (ICMR) beamformer provides a flexible way to control the specified amount of the interference attenuation and noise suppression; but it requires accurate estimation of the manifold vector of the interference sources, which is challenging to achieve in real-world applications. To address this issue, we introduce an interference-controlled maximum noise reduction network (ICMRNet) in this study, which is a deep neural network (DNN)-based method for manifold vector estimation. With densely connected modified conformer blocks and the end-to-end training strategy, the interference manifold is learned directly from the observation signals. This approach, akin to ICMR, adeptly adapts to time-varying interference and demonstrates superior convergence rate and extraction efficacy as compared to the linearly constrained minimum variance (LCMV)-based neural beamformers when appropriate attenuation factors are selected. Moreover, via learning-based extraction, ICMRNet effectively suppresses reverberation components within the target signal. Comparative analysis against baseline methods validates the efficacy of the proposed method.
Yichen Yang 0010, Ningning Pan, Wen Zhang 0002, Chao Pan 0001, Jacob Benesty, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Design of Maximum Directivity Beamformers With Linear Acoustic Vector Sensor Arrays
abstract
This paper studies the design of maximum directivity factor (MDF) beamformers based on uniform linear arrays (ULAs) consisting of acoustic vector sensors (AVSs). We first derive the main lobe constraints, which ensure that the beamformer's beampattern achieves a maximum in the look direction, and prove that any beamformer that satisfies the proposed constraints can be written as the sum of two orthogonal beamformers: the maximum white noise gain (MWNG) beamformer and a reduced-rank beamformer. Then, we derive the MDF beamformer by maximizing the directivity factor (DF) under the deduced constraints. We also derive a robust version of the MDF beamformer, which can keep the WNG above a pre-specified level. Compared to the conventional MDF beamformer based on ULAs with omnidirectional microphones, the designed MDF beamformer with uniform linear AVS arrays (ULAVSAs) can steer the beampattern to any look direction in the 3-dimensional space and achieves a higher directivity. The proposed MDF beamformer also outperforms the two-step MDF beamformer with ULAVSAs since it maximizes the DF. The proposed methods are validated through simulations as well as real experiments.
Xueqin Luo, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty, Wen Zhang 0002, Mengyao Zhu 0003, Chunjian Li
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 CGMM-Based Sound Zone Generation Using Robust Pressure Matching With ATF Perturbation Constraints
abstract
Personal sound zone (PSZ) refers to the technique that uses an array of loudspeakers and digital signal processing tools to achieve spatial soundfield control. To generate the target sound zones, this technique generally requires to know the acoustic transfer functions (ATFs) between the loudspeakers and the spots where soundfields are to be controlled. In practical applications, however, the true ATFs are never accessible and they have to be measured or estimated. Due to many sophisticated reasons, the measured ATFs generally deviate from the true ones, which may lead to significant degradation in performance of sound zone reproduction. In this work, a robust pressure matching (RPM) algorithm is presented for sound zone generation. It exploits a complex Gaussian mixture model (CGMM) to model the ATFs and their perturbations. The CGMM parameters are estimated using the expectation-maximization (EM) algorithm. To improve the robustness of the pressure matching method, an uncertainty constraint is applied to the ATF estimates and the pressure matching problem is then formulated as one of biconvex optimization. The coordinate descent algorithm is subsequently used to solve the optimization problem, thereby obtaining the optimal control filter. In comparison with the existing pressure matching methods without considering the effect of ATF perturbations, the presented algorithm is able to achieve lower normalized signal distortion energy and higher signal to interference ratio. Numerical simulations justify the effectiveness of the presented algorithm as well as its advantages over the traditional methods.
Junqing Zhang, Liming Shi, Mads Græsbøll Christensen, Wen Zhang 0002, Lijun Zhang 0004, Jingdong Chen
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Robust Pressure Matching with ATF Perturbation Constraints for Sound Field Control
abstract
Sound field control systems deployed in room acoustic environments require knowing the acoustic channel impulse responses between the loudspeakers and matching microphones, which are challenging to estimate accurately due to perturbations caused by such factors as temperature changes and sensors’ position mismatches. To deal with this issue, a robust pressure matching algorithm is developed in this work where a perturbation term of the acoustic transfer function (ATF) is modeled as a Gaussian process, based on which an uncertainty constraint is applied to limit the impact of perturbation on pressure matching. This constrained problem is formulated as one of biconvex optimization, and a coordinate descent algorithm is adopted to estimate the optimal control filter. Simulations are performed and results show that the proposed method is able to achieve more accurate control as compared to the standard pressure matching algorithm in the presence of ATF perturbations.
Junqing Zhang, Liming Shi, Mads Græsbøll Christensen, Wen Zhang 0002, Lijun Zhang 0004, Jingdong Chen
ICASSP4
2021 Wave-domain active noise control over distributed networks of multi-channel nodes
Yuchen Dong, Jie Chen 0022, Wen Zhang 0002
Signal Process.3
2021 Multiple Circular Arrays of Vector Sensors for Real-Time Sound Field Analysis
abstract
This article proposes multiple circular arrays of vector sensors for analyzing the three dimensional sound field. By exploiting the fact that a finite number of spatial basis functions can represent the sound field within a region, the designed arrays allow the analysis of the sound field around the arrays sequentially by either a frequency-domain method or a time-domain method. The arrays are compact and thus are suitable for integrating with small-sized devices. The time-domain method is of low latency and computationally simple, and thus are suitable for providing the devices with spatial acoustic features in real-time. The effectiveness of the proposed arrays and the sound field analysis methods are confirmed by simulations, and compared with another array.
Fei Ma 0008, Thushara D. Abhayapala, Wen Zhang 0002
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Spatial Active Noise Control in Rooms Using Higher Order Sources
abstract
All spatial active noise control (ANC) systems, when deployed in typical room environments, have time-varying acoustic channels between the secondary sources and the error microphones. The conventional online secondary path modeling techniques, which introduces additive auxiliary random noise to estimate the secondary paths, become challenging especially in a multichannel setup. In this work, we propose to use higher-order variable-directivity sound sources as secondary sources for spatial ANC, in which both the interior residual noise field within the control region and exterior sound field due to secondary source radiation are jointly controlled. The aim of controlling the exterior sound field is to minimize room reverberation generated by the secondary sources so that the secondary paths in the proposed algorithm can be approximated as free-field propagation and thus can be pre-calibrated. The system is implemented in an adaptive manner to track noise variations. The results show that the proposed method can effectively cancel spatial noise field and control exterior sound field at an acceptable low level in time-varying room environments.
Junqing Zhang, Wen Zhang 0002, Jihui Zhang 0006, Thushara D. Abhayapala, Lijun Zhang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Distributed Wave-Domain Active Noise Control Based on the Diffusion Strategy
abstract
Conducting the spatial active noise control (ANC) in wave-domain has been shown advantageous over conventional point-based methods. In the existing schemes, signals at all error microphones are collected and processed in a centralized manner to update the secondary source driving signals. The high computational complexity of this centralized strategy represents one of the major challenges for applying multi-channel ANC systems in large-scale applications. In order to address this issue, this work presents a distributed wave-domain ANC scheme by resorting to distributed optimization techniques. The global ANC problem is formulated as a sum of local costs, and the diffusion adaptation strategy is subsequently utilized to provide a distributed solution that only requires local information exchanges. Simulation results show that the proposed algorithm achieves sufficiently good performance compared to its centralized counterpart.
Yuchen Dong, Jie Chen 0022, Wen Zhang 0002
ICASSP3
2020 Distributed Wave-Domain Active Noise Control Based on the Diffusion Adaptation
abstract
Conducting the spatial active noise control (ANC) in wave-domain has been shown advantageous over conventional point-based methods. In the existing schemes, signals at all error microphones are collected and processed in a centralized manner to update the secondary source driving signals. The high computational complexity of this centralized strategy represents one of the major challenges for applying multi-channel ANC systems in large-scale applications. In order to address this issue, this work presents a distributed wave-domain ANC scheme by resorting to distributed optimization techniques. The global ANC problem is formulated as a sum of local costs, and the diffusion adaptation strategies are subsequently utilized to provide a distributed solution that only requires local information exchanges. Simulation results show that the proposed algorithms achieve sufficiently good performance compared to its centralized counterpart.
Yuchen Dong, Jie Chen 0022, Wen Zhang 0002
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Active Control of Outgoing Broadband Noise Fields in Rooms
abstract
Active noise control system has been actively researched over the past half century, and implemented to reduce noises in ducts, headsets, and inside several automobile models. However, active control of noise fields, and specifically broadband noise fields, in rooms is still a barely explored topic due to the difficult of obtaining the reference signals, the challenge of online secondary path estimation, and the causal control constraint. In this paper, an active noise control system is developed that can cancel outgoing broadband noise fields in rooms. The proposed system decomposes the noise field on a sphere surrounding the noise sources into spherical harmonic modes, exploiting their directionality to generate the reference signals and remove the need for online secondary path estimation. A time-domain sound field separation algorithm and a time-wave domain adaptive algorithm allow the proposed system to meet the causal control constraint. Simulation results demonstrate that the proposed system can cancel broadband noise field globally in a room without suffering from the secondary source feedback problem.
Fei Ma 0008, Wen Zhang 0002, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Robust Sparse Multichannel Active Noise Control
abstract
Multichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ1-norm constraint to the standard MC-ANC helps to reduce the complexity of the system and accelerate the convergence rate. However, the convergence performance of ℓ1-norm constrained MC-ANC (cℓ1-MC-ANC) degrades significantly in reverberant environments. In this paper, we analyze the necessity of using sparsity-inducing algorithms with distinct zero-attracting strengths over loudspeakers, and then derive three algorithms of this kind in the complex domain. Simulation results show that, compared to cℓ1-MC-ANC, the proposed algorithms exhibit faster convergence or higher noise reduction at steady state in both free field and reverberant environments.
Jingli Xie, Danqi Jin, Wen Zhang 0002, Xiao-Lei Zhang 0001, Jie Chen 0022, DeLiang Wang
ICASSP3
2019 2.5D Multizone Reproduction with Active Control of Scattered Sound Fields
abstract
Multizone reproduction has been focused on reproducing sounds in an empty listening space. However, there are always scatterers such as human heads in sound zones, generating scattered sound fields and causing degraded system performance. In this work, we develop a modal-domain method for 2.5D multizone reproduction with a solid object in the bright zone. Analytical expressions of the incident and scattered fields are developed. We then propose an active control strategy to correct the scattering effect. In the reproduction stage, we use the weighted mode matching approach to achieve the optimal control over the entire region. Simulation results show that in comparison with the conventional method which does not consider the scattering effect, the proposed method can achieve higher acoustic contrast performance over a broadband frequency range.
Junqing Zhang, Wen Zhang 0002, Thushara D. Abhayapala, Jingli Xie, Lijun Zhang 0004
ICASSP2
2018 Reference Signal Generation for Broadband ANC Systems in Reverberant Rooms
abstract
One major issue of implementing broadband active noise control systems in reverberant rooms is the lack of reference signals. In this work, by exploiting the spatial sound field characteristics, a time-domain sound field separation method is developed to generate the reference signal for broadband active noise control systems in reverberant rooms. The time-domain sound field separation method separates the outgoing field produced by the primary source from the secondary source feedback and room reverberation on a spherical array, based on spherical harmonic decomposition of the sound field and the radial particle velocity measured by the array. Both the effectiveness of the time-domain sound field separation method and the applicability of the separated outgoing field as the reference signal for a broadband active noise control system in a reverberant room are demonstrated by simulations.
Fei Ma 0008, Wen Zhang 0002, Thushara D. Abhayapala
ICASSP2
2018 2.5D Multizone Reproduction Using Weighted Mode Matching
abstract
The mode matching based multizone reproduction has mainly been focused on a purely 2D theory which is inadequate to fit the 3D reality. Its extension to the 3D theory however requires many secondary sources and a high computational complexity. In this paper, a weighted mode matching approach is developed for 2.5D multizone reproduction. The multizone soundfield is reproduced in the horizontal plane within a circular control region using the loudspeakers modelled as 3D point sources. We propose weighting the Bessel-spherical harmonic modes for 2.5D reproduction and a matching between the desired and reproduced soundfields over the entire control region. Simulation results show that in comparison with the conventional 2.5D reproduction method a more accurate reproduction is achieved using the proposed weighting approach.
Wen Zhang 0002, Junqing Zhang, Thushara D. Abhayapala, Lijun Zhang 0004
ICASSP1
2018 Active Noise Control Over Space: A Wave Domain Approach
abstract
Noise control and cancellation over a spatial region is a fundamental problem in acoustic signal processing. In this paper, we utilize wave-domain adaptive algorithms to iteratively calculate the secondary source driving signals and to cancel the primary noise field over the control region. We propose wave-domain active noise control algorithms based on two minimization problems: first, minimizing the wave-domain residual signal coefficients, and second, minimizing the acoustic potential energy over the region, and derive the update equations with respect to two variables, the loudspeaker weights and wave-domain secondary source coefficients. Simulation results demonstrate the effectiveness of the proposed algorithms, more specifically the convergence speed and the noise cancellation performance in terms of the noise reduction level and acoustic potential energy reduction level over the entire spatial region.
Jihui Zhang 0006, Thushara D. Abhayapala, Wen Zhang 0002, Prasanga N. Samarasinghe, Shouda Jiang
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Spatial Noise-Field Control With Online Secondary Path Modeling: A Wave-Domain Approach
abstract
Due to strong interchannel interference in multichannel active noise control (ANC), there are fundamental problems associated with the filter adaptation and online secondary path modeling remains a major challenge. This paper proposes a wave-domain adaptation algorithm for multichannel ANC with online secondary path modelling to cancel tonal noise over an extended region of two-dimensional plane in a reverberant room. The design is based on exploiting the diagonal-dominance property of the secondary path in the wave domain. The proposed wave-domain secondary path model is applicable to both concentric and nonconcentric circular loudspeakers and microphone array placement, and is also robust against array positioning errors. Normalized least mean squares-type algorithms are adopted for adaptive feedback control. Computational complexity is analyzed and compared with the conventional time-domain and frequency-domain multichannel ANCs. Through simulation-based verification in comparison with existing methods, the proposed algorithm demonstrates more efficient adaptation with low-level auxiliary noise.
Wen Zhang 0002, Christian Hofmann 0001, Michael Buerger, Thushara D. Abhayapala, Walter Kellermann
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Online secondary path modelling in wave-domain active noise control
abstract
The performance of an ANC system largely depends on the availability of an accurate secondary path model. This is however a major challenge in multichannel ANC where the computational complexity increases significantly with the number of secondary sources and error sensors. This paper proposes wave-domain adaptive processing algorithm for multichannel ANC with online secondary modelling to cancel a tonal noise over a region of space within a reverberant room. The design is based on exploiting a special property of the secondary path model in the wave domain. A feedback control system is implemented, where a single microphone array is placed at the boundary of the control region to measure the residual signals, and a loudspeaker array reproduces secondary sources to generate the anti-noise signals and auxiliary noise for secondary path modelling. Through experimental verification in comparison with existing methods the proposed algorithm demonstrates more efficient adaptation with low-level auxiliary noise.
Wen Zhang 0002, Christian Hofmann 0001, Michael Buerger, Thushara D. Abhayapala, Walter Kellermann
ICASSP1
2017 Direct-to-Reverberant Energy Ratio Estimation Using a First-Order Microphone
abstract
The direct-to-reverberant ratio (DRR) is an important characterization of a reverberant environment. This paper presents a novel blind DRR estimation method based on the coherence function between the sound pressure and particle velocity at a point. First, a general expression of coherence function and DRR is derived in the spherical harmonic domain, without imposing assumptions on the reverberation. In this paper, DRR is expressed in terms of the coherence function as well as two parameters that are related to statistical characteristics of the reverberant environment. Then, a method to estimate the values of these two parameters using a microphone system capable of capturing first-order spherical harmonics is proposed, under three assumptions which are more realistic than the diffuse field model. Furthermore, a theoretical analysis on the use of plane wave model for direct path signal and its effect on DRR estimation is presented, and a rule of thumb is provided for determining whether the point source model should be used for the direct path signal. Finally, the ACE challenge dataset is used to validate the proposed DRR estimation method. The results show that the average full band estimation error is within 2 dB, with no clear trend of bias.
Hanchi Chen, Thushara D. Abhayapala, Prasanga N. Samarasinghe, Wen Zhang 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Spatial feature learning for robust binaural sound source localization using a composite feature vector
abstract
The performance of binaural speech source localization systems can be significantly impacted by an imperfect selection of spatial localization cues, due to the limited bandwidth of speech, and the effects of noise. In order to mitigate these impacts, this paper presents a novel method that combines a deterministic localization approach with a spatial feature learning process. Here, we (i) obtain a composite feature vector derived from analysing the mutual information between different spatial cues and (ii) estimate the optimum feature combination that minimizes the angular localization error in three dimensional space. The performance of the proposed mutual information based feature learning approach is evaluated for a range of speech inputs and noise conditions. We also demonstrate that the proposed approach improves the localization accuracy and its robustness, with respect to traditional localization algorithms, especially in the relatively low signal-to-noise ratio localization scenarios.
Dumidu S. Talagala, Wen Zhang 0002, Thushara D. Abhayapala
ICASSP3
2016 Sparse complex FxLMS for active noise cancellation over spatial regions
abstract
In this paper, we investigate active noise control over large 2D spatial regions when the noise source is sparsely distributed. The l1relaxation technique originated from compressive sensing is adopted and based on that we develop the algorithm for two cases: multipoint noise cancellation and wave domain noise cancellation. This results in two new variants (i) zero-attracting multi-point complex FxLMS and (ii) zero-attracting wave domain complex FxLMS. Both approaches use a feedback control system, where a microphone array is distributed over the boundary of the control region to measure the residual noise signals and a loudspeaker array is placed outside the microphone array to generate the anti-noise signals. Simulation results demonstrate the performance and advantages of the proposed methods in terms of convergence rate and spatial noise reduction levels.
Jihui Zhang 0006, Thushara D. Abhayapala, Prasanga N. Samarasinghe, Wen Zhang 0002, Shouda Jiang
ICASSP4
2015 Binaural localization of speech sources in 3-D using a composite feature vector of the HRTF
abstract
Binaural localization of speech sources in 3-D, using head-related transfer functions (HRTFs), always suffers elevation ambiguity due to the limited high frequency spectral information available at the receivers. This paper presents a method that overcomes this limitation by exploiting the interaural phase and magnitude features present in the HRTF. We (i) introduce a new feature vector that combines these two sets of features in a non-linear fashion, and (ii) propose a mechanism to extract this feature vector free from distortion by the speech spectra. The performance of the proposed method is evaluated and compared with a correlation-based HRTF database matching approach and a two-step localization technique for multiple source positions, HRTFs (individuals) and speech inputs. The results suggest that up to 20% improvement in localization performance can be achieved for moderate signal-to-noise ratios.
Dumidu S. Talagala, Wen Zhang 0002, Thushara D. Abhayapala
ICASSP3
2014 Three Dimensional Sound Field Reproduction using Multiple Circular Loudspeaker Arrays: Functional Analysis Guided Approach
abstract
Three dimensional sound field reproduction based on higher order Ambisonics requires the placement of loudspeakers on a sphere that surrounds the target reproduction region. The deployment of a spherical array is not trivial especially for implementation in real rooms where the placement flexibility is highly desirable. This paper proposes a design of multiple circular loudspeaker arrays for reproducing three dimensional sound fields originating from a limited region of interest. We apply a functional analysis framework to formulate the sound field reproduction problem in a closed form. Secondary source distributions and target sound fields are modeled as two Hilbert spaces and mapped by an integral operator and its adjoint operator, from which a self-adjoint operator is constructed and the singular value decomposition is applied to represent source distributions and sound fields with two sets of interrelated singular functions. We derive the solutions for a circular secondary source arrangement and propose the design of placing multiple circular loudspeaker arrays only over the limited region of interest. Such a design allows for non-spherical and non-uniform loudspeaker placement and thus provides a flexible array arrangement. The reproduction accuracy of the proposed method is verified through numerical simulations.
Wen Zhang 0002, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Efficient Multi-Channel Adaptive Room Compensation for Spatial Soundfield Reproduction Using a Modal Decomposition
abstract
Mitigating the effects of reverberation is a significant challenge for real-world spatial soundfield reproduction, but the necessity of a large number of reproduction channels increases the complexity and presents several challenges to existing listening room compensation techniques. In this paper, we present an adaptive room compensation method to overcome the effects of reverberation within a region, using a model description of the reverberant soundfield. We propose the reverberant channel estimation and compensation be carried out in a single step using completely decoupled adaptive filters; thus, reducing the complexity of the overall process. We compare the soundfield reproduction performance with existing adaptive and nonadaptive room compensation methods through several simulation examples. The performance of the proposed method is comparable to existing techniques, and achieves a normalized wideband region reproduction error of 1% at a signal-to-noise ratio of 50 dB, within a 1 m radius region of interest using 60 loudspeakers and 55 microphones at frequencies below 1 kHz. Robust behavior of the room compensator is demonstrated down to direct-to-reverberant-path power ratios of -5 dB. Overall, the results suggest that the proposed method can diagonalize the room compensation system, leading to a more robust and parallel implementation for spatial soundfield reproduction.
Dumidu S. Talagala, Wen Zhang 0002, Thushara D. Abhayapala
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Active acoustic echo cancellation in spatial soundfield reproduction
abstract
The equalization of reverberation effects is essential for spatial soundfield reproduction, but estimation of the reverberant channel presents several challenges to existing equalization techniques. This paper presents a method of active acoustic echo cancellation (AEC) for soundfield reproduction applications, using a modal description of the reverberant soundfield. We describe how individual modes of the measured soundfield can be equalized adaptively, thus reducing the complexity of the channel estimation process. AEC and reproduction performance is compared with existing adaptive and non-adaptive equalization techniques through simulation examples. Equalization performance is comparable to existing methods, achieving a normalized region reproduction error of 1% and echo return loss enhancement of 15 - 30 dB at 50 dB SNR. The results suggest that the proposed model can be used to obtain a parallel implementation of a room equalizer for active AEC.
Dumidu S. Talagala, Wen Zhang 0002, Thushara D. Abhayapala
ICASSP2
2013 Broadband DOA Estimation Using Sensor Arrays on Complex-Shaped Rigid Bodies
abstract
Sensor arrays mounted on complex-shaped rigid bodies are a common feature in many practical broadband direction of arrival (DOA) estimation applications. The scattering and reflections caused by these rigid bodies introduce complexity and diversity in the frequency domain of the channel transfer function, which presents several challenges to existing broadband DOA estimators. This paper presents a novel high resolution broadband DOA estimation technique based on signal subspace decomposition. We describe how broadband signals can be decomposed into narrow subband components, and combined such that the frequency domain diversity is retained. The DOA estimation performance is compared with existing techniques using a uniform circular array and a sensor array on a hypothetical rigid body. An improvement in closely spaced source resolution of up to 6 dB is observed for the sensor array on the hypothetical rigid body, in comparison to the uniform circular array. The results suggest that frequency domain diversity, introduced by complex-shaped rigid bodies, can provide higher resolution and clearer separation of closely spaced broadband sound sources.
Dumidu S. Talagala, Wen Zhang 0002, Thushara D. Abhayapala
IEEE Trans. Speech Audio Process.2
2012 On High-Resolution Head-Related Transfer Function Measurements: An Efficient Sampling Scheme
abstract
This paper deals with two important questions associated with HRTF measurement: 1) “what is the required angular resolution?,” and 2) “what is the most suitable sampling scheme?.” The paper shows that a well-defined finite number of spherical harmonics can capture the head-related transfer function (HRTF) spatial variations in sufficient detail, which is defined as the HRTF spatial dimensionality. For the 20-kHz audible frequency range, the value of the dimensionality means a high-directional resolution HRTF measurement is required. Considering such a high-resolution measurement, a number of sampling criteria have been identified from both mechanical setup and data processing aspects. Different sampling candidates are then compared to demonstrate that the best method which satisfies all requirements is the class termed as IGLOO. A fast spherical harmonic transform algorithm based on the IGLOO scheme is developed to accelerate the high-resolution data analysis. The proposed method is validated through simulation and experimental data acquired from a KEMAR mannequin.
Wen Zhang 0002, Mengqiu Zhang, Rodney A. Kennedy, Thushara D. Abhayapala
IEEE Trans. Speech Audio Process.1
2009 Modal expansion of HRTFs: Continuous representation in frequency-range-angle
abstract
This paper proposes a continuous HRTF representation in both 3D spatial and frequency domains. The method is based on the acoustic reciprocity principle and a modal expansion of the wave equation solution to represent the HRTF variations with different variables in separate basis functions. The derived spatial basis modes can achieve HRTF near-field and far-field representation in one formulation. The HRTF frequency components are expanded using Fourier Spherical Bessel series for compact representation. The proposed model can be used to reconstruct HRTFs at any arbitrary position in space and at any frequency point from a finite number of measurements. Analytical simulated and measured HRTFs from a KEMAR are used to validate the model.
Wen Zhang 0002, Thushara D. Abhayapala, Rodney A. Kennedy, Ramani Duraiswami
ICASSP1
2009 Efficient Continuous HRTF Model Using Data Independent Basis Functions: Experimentally Guided Approach
abstract
This paper introduces a continuous functional model for head-related transfer functions (HRTFs) in the horizontal auditory scene. The approach uses a separable representation consisting of a Fourier-Bessel series expansion for the spectral components and a conventional Fourier series expansion for the spatial components. Being independent of the data, these two sets of basis functions remain unchanged for all subjects and measurement setups. Hence, the model can transform an individualized HRTF to a subject specific set of coefficients. A continuous functional model is also developed in the time domain. We show the efficient model performance in approximating experimental measurements by using the HRTF measurements from a KEMAR manikin and the synthetic data from the spherical head model. The statistical results are determined from a 50-subject HRTF data set. We also corroborate the predictive capability of the proposed model. The model has near optimal performance, which can be ascertained by comparison with the standard principle component analysis (PCA) and discrete Karhunen-Loeve expansion (KLE) methods at the measurement points and for a given number of parameters.
Wen Zhang 0002, Rodney A. Kennedy, Thushara D. Abhayapala
IEEE Trans. Speech Audio Process.1
2008 Iterative extrapolation algorithm for data reconstruction over sphere
abstract
Given limited or incomplete measurement data on a sphere, a new iterative algorithm is proposed on how to extrapolate signal over the whole sphere. The algorithm is based on a priori assumption that the Fourier decomposition of the signal on the sphere has finite degree of spherical harmonic coefficients, that is, the signal is mode limited. The algorithm is a simple iteration involving only the spherical harmonic decomposition. It is proven that the algorithm converges to the original signal over observation region and the convergence rate is lower bounded by the largest eigenvalue of an associated Fredholm integral equation.
Wen Zhang 0002, Rodney A. Kennedy, Thushara D. Abhayapala
ICASSP1
2006 UWB Spatia - Frequency Channel Characterization
abstract
This paper investigates the spatial-frequency channel characterization of ultra-wideband (UWB) wireless communication systems. First, a novel frequency dependent UWB channel model is constructed based on the theory of electromagnetic diffraction mechanism, which causes the field strength to vary with the frequency in each multipath. Then, we build a space-frequency model, which includes spatial characteristics such as angular power spectrum, and physical sampling points in space. The space-frequency model has two special cases (i) discrete multipath model, and (ii) cluster model, which can be readily used to generate channel data for any arbitrary set of sensor locations. The reconstruction results from channel measurements show the accurateness of the novel frequency dependent model, with reconstruction error decreasing by 40%, compared to the traditional Turin model
Wen Zhang 0002, Thushara D. Abhayapala, Jian (Andrew) Zhang
VTC Spring1