VLDB 2026 Research / reviewers in the wild / expert
Jacob Benesty
dblp:47/2472
· DBLP profile ↗
283ranked-venue papers
39as first author
72since 2021 · last 2026
0000-0002-0036-5865ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 187 · 29 first-author · 50 since 2021Artificial intelligence and machine learning · 95 · 10 first-author · 23 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing the NP-VSS-NLMS algorithm based on insights from its stochastic model
Augusto Cesar Becker, Eduardo Vinicius Kuhn, Jacob Benesty, Rui Seara |
Signal Process. | 3 |
| 2026 | A unified Bayesian perspective for conventional and robust adaptive filters
Leszek Szczecinski, Jacob Benesty, Eduardo Vinicius Kuhn |
Signal Process. | 2 |
| 2026 | On adaptive multichannel dereverberation based on dichotomous coordinate descent and data-reuse techniques
Wenxing Yang, Jilu Jin, Jingdong Chen, Jacob Benesty |
Signal Process. | 5 |
| 2026 | Design of Low-Rank differential beamformers with constrained directivity or robustness
Kunlong Zhao, Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
Signal Process. | 6 |
| 2026 | A third-order tensor decomposition based algorithm for speech dereverberation
Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
Signal Process. | 5 |
| 2026 | A Phase-Based Feature for Gas Detection Under Unstable Preheating Condition for E-NosesabstractElectronic noses (e-noses) utilizing metal oxide semiconductor (MOS) gas sensors are widely used; however, they typically require several days of preheating after power-on, whether during initial startup or following a power interruption, to reach a stable operation. Such prolonged preheating significantly limits immediate deployment, particularly in portable systems, and has rarely been systematically addressed in the existing literature. To tackle this challenge, we propose a phase-based feature extraction method to capture stable signal patterns under temperature modulation to address baseline drift during the unstable preheating stage. Building on this feature, we develop a detection method that markedly outperforms conventional magnitude-based techniques, which typically exhibit very low detection probabilities during early preheating. The proposed method achieves high-precision gas detection within just 14 hours, and reaches even better detection probability in 2 hours that conventional methods require 154 hours to match when most commercial sensors remain in the pre-conditioning stage. By incorporating the phase-based feature extraction strategy, the required stabilization time is substantially shortened, enabling rapid and reliable gas detection, and facilitating real-time deployment of e-nose systems in critical applications such as emergency response and industrial safety. Lihua Guo, Zhengqiao Zhao, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 4 |
| 2026 | A Nonparametric Variable Forgetting Factor Recursive Least-Squares AlgorithmabstractThe forgetting factor is the main control parameter of the recursive least-squares (RLS) algorithm, which is set to balance between the estimate accuracy and the tracking capability. Targeting a better performance compromise, different variable forgetting factor RLS (VFF-RLS) algorithms have been previously developed. Usually, they require additional control mechanisms and/or extra parameters, which are difficult to handle in practice. In this letter, we present an elegant yet practical VFF-RLS algorithm, in the framework of system identification. Simulation results obtained in the context of acoustic echo cancellation indicate its reliable performance. Constantin Paleologu, Jacob Benesty, Radu-Andrei Otopeleanu, Silviu Ciochina |
IEEE Signal Process. Lett. | 2 |
| 2025 | Advances in Microphone Array Processing and Multichannel Speech EnhancementabstractThis paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas. The paper examines foundational developments in microphone array design and optimization, showcasing innovations that improved sound acquisition and enhanced speech intelligibility in noisy and reverberant environments. It then introduces recent advancements and cutting-edge research in the field, particularly the integration of deep learning techniques such as all-neural beamformers. The paper also explores critical applications, discussing their evolution and current state-of-the-art technologies that significantly impact user experience. Finally, the paper outlines future research directions, identifying challenges and potential solutions that could drive further innovation in these fields. By providing a comprehensive overview and forward-looking perspective, this paper aims to inspire ongoing research and contribute to the sustained growth and development of microphone arrays and multichannel speech enhancement. Gongping Huang, Jesper Rindom Jensen, Jingdong Chen, Jacob Benesty, Mads Græsbøll Christensen, Akihiko Sugiyama, Gary W. Elko, Tomas Gänsler |
ICASSP | 4 |
| 2025 | Microphone Array Beamforming for Speech Enhancement Based on Dynamic Mode DecompositionabstractMicrophone array beamforming is widely used to extract desired speech signals from noisy environments. While most research in this area focuses on utilizing spatial information, less attention is given to the intrinsic physical mechanisms underlying microphone array observations. This paper aims to address this gap by exploring these underlying factors through dynamic mode decomposition (DMD). Our contributions are twofold. 1) We develop a DMD-based signal model for microphone arrays to capture the relationships between observation signals at adjacent microphones. 2) We introduce a DMD-based preprocessing method and a corresponding beamforming approach based on this model. Simulation results show that our proposed method significantly enhances performance compared to conventional beamforming techniques. Wei Liu 0177, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
ICASSP | 6 |
| 2025 | On the Design of Low-Rank Differential Beamformers with Nonuniform Linear Microphone ArraysabstractKronecker product beamforming is an effective technique for designing beamformers with nonuniform linear arrays (NULAs). However, current techniques are restricted to NULAs with specific configurations, where the steering vector of the array is represented as a Kronecker product of steering vectors from smaller virtual arrays. This paper overcomes these constraints by proposing a novel approach to designing Kronecker product beamformers for NULAs from a low-rank perspective. Our approach involves decomposing the NULA into overlapping subarrays and organizing the sensor signals from these subarrays into a matrix. We then apply filters to both sides of this matrix to produce an output, which is then converted into a low-rank beamforming process. This method is highly adaptable and can be utilized for NULAs with any number of microphones. Hanchen Pei, Gongping Huang, Jilu Jin, Jacob Benesty, Jingdong Chen |
ICASSP | 5 |
| 2025 | Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech EnhancementabstractSuperdirective beamformers are highly effective at suppressing directional interference and diffuse noise, but their practical use is often constrained by the problem of white noise amplification. Robust superdirective beamforming methods typically address this by imposing a constraint on the white noise gain (WNG). However, determining the appropriate WNG threshold in varying noise environments remains unclear. This paper introduces a data-driven approach to estimating the optimal WNG threshold. Subsequently, a more versatile and robust superdirective beamformer is developed by solving a quadratic eigenvalue problem (QEP). Experimental results show that this method outperforms traditional superdirective beamformers, which rely on a WNG threshold set through a fixed search range. Importantly, this approach functions as a distortionless beamformer, maintaining high fidelity of the desired acoustic signal and allowing for additional post-filtering if required. Hanchen Pei, Gongping Huang, Jilu Jin, Zhizheng Wu 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 7 |
| 2025 | Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech EnhancementabstractSuperdirective beamformers, used with small microphone arrays, are highly attractive due to their high directivity and frequency-invariant beampatterns, making them well-suited for processing broadband acoustic and speech signals. However, these beamformers are very sensitive to array imperfections such as sensor mismatches and self-noise. To improve robustness, robust superdirective (RSD) beamformers have been developed, employing techniques such as diagonal loading or white-noise-gain constraints during their derivation. Although RSD beamformers offer enhanced robustness compared to classical superdirective beamformers, they cannot achieve the maximum directivity factor and lose some frequency-invariant properties, resulting in a beamwidth that is wider at low frequencies and narrower at high frequencies. As a result, RSD beamformers do not fully meet the criteria of true superdirective beamformers, providing less effective noise reduction and introducing some speech distortion. Post-filtering methods have been developed to improve noise reduction after RSD beamforming, but they often fail to address the distortion issues, especially when the speech source deviates from the array’s look direction. To overcome this limitation, this paper proposes a joint optimization approach that combines post-filtering with RSD beamformers. By using the output of RSD beamformers as input data and considering various deviations in look directions and array mismatches, we train a post-filtering network to further enhance the beamformer’s output. Experimental results on speech enhancement demonstrate the effectiveness and robustness of the proposed method. Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2025 | DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone ArraysabstractDirection-of-arrival (DOA) estimation is a key process in microphone array systems. The steered response power-based minimum variance distortionless response (SRP-MVDR) method performs very well in challenging acoustic environments but suffers from exponential complexity as the number of microphones increases. To improve the efficiency of SRP-MVDR for real-time applications, we propose a Kronecker product-based SRP-MVDR (SRP-KPMVDR) method designed for large rectangular microphone arrays. This approach begins with a rank-one approximation that represents the signal covariance matrix of a rectangular microphone array in Kronecker product form, which is essential for SRP-MVDR estimation. By utilizing the Kronecker product properties, the complex matrix inversion in SRP-MVDR is simplified to the inversion of two smaller matrices, significantly reducing computational complexity. Simulation results show that the SRP-KPMVDR method achieves comparable performance to the traditional SRP-MVDR while greatly decreasing the computational demands. Yichen Zeng, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2025 | Radiation and Directivity Analysis of a Vibrating Dome-Shaped Radiator Mounted on an Infinite BaffleabstractAccurate modeling and analysis of a radiator mounted on an infinite baffle are crucial for understanding its acoustic radiation characteristics. This paper investigates the radiation behavior of a convex dome-shaped radiator in such a condition, showing that, under the far-field approximation, the pressure field is the three-dimensional Fourier transform of the axisymmetric surface velocity distribution. The study includes a comparison of various typical velocity distributions in terms of their directivity factor (DF) and radiated sound power. Current velocity distributions often suffer from nulls in the DF, so we provide a detailed analysis to uncover the causes of this issue. To address this, we propose an equalization filter designed to smooth the DF across the entire frequency range. Simulations are performed to validate the theoretical findings and to showcase the improved performance of the proposed approach. Junqing Zhang, Wen Zhang 0002, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2025 | Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar GeometryabstractDifferential microphone arrays (DMAs) have garnered significant attention in recent research and development due to their high directivity and frequency-invariant beampatterns. However, DMAs frequently encounter substantial white noise amplification, which limits their practical applications. This paper addresses this issue by introducing a general method for designing robust DMAs with microphone arrays of arbitrary planar topology. The proposed approach approximates the beampattern using the Jacobi-Anger series expansion and constrains the white noise gain (WNG) to a specified value. This minimizes the error between the beampattern and the ideal directivity pattern while ensuring a reasonable level of robustness. A closed-form solution for the robust differential beamformer filter is derived using the quadratic eigenvalue problem (QEP) method. Simulation results demonstrate the feasibility and effectiveness of the proposed approach. Kunlong Zhao, Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 6 |
| 2025 | On the Design of a Robust Superdirective Beamformer and Topology Parameter Optimization with Frustum-Shaped Microphone Arrays Featuring Multiple Rings
Kunlong Zhao, Gongping Huang, Jingdong Chen, Jacob Benesty, Zoran Cvetkovic |
INTERSPEECH | 5 |
| 2025 | A MISO acoustic echo cancellation algorithm based on a two-layer filter decomposition
Zi Cao, Tianci Yan, Jingdong Chen, Jacob Benesty |
Signal Process. | 5 |
| 2025 | Automatic regularization for linear MMSE filters
Daniel Gomes de Pinho Zanco, Leszek Szczecinski, Jacob Benesty |
Signal Process. | 3 |
| 2025 | An update rule for multiple source variances estimation using microphone arrays
Fan Zhang 0001, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
Speech Commun. | 4 |
| 2025 | MPPCAD: Minimum Power Pattern Constrained Adaptive Differential BeamformingabstractThis paper investigates the design of adaptive differential beamforming using small-spacing linear microphone arrays. We express the differential beamformer as a linear function of the target beampattern coefficients through orthogonal polynomial expansions. Consequently, the design of the beamformer reduces to optimizing these coefficients. To ensure that the maximum array response consistently aligns with the look direction, we derive constraints on the target beampattern coefficients, resulting in two convex sets for the first two orders of beampatterns. This approach uncovers numerous effective beampatterns beyond traditional options such as dipole, cardioid, supercardioid, and hypercardioid. By minimizing the power of the array output while adhering to the beampattern constraints, we develop the MPPCAD beamformer. Simulation results demonstrate that the proposed beamformer significantly enhances speech quality compared to classical differential beamformers. Fan Zhang 0001, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 4 |
| 2024 | Beamforming Through Online Convex Combination of Differential BeamformersabstractThanks to their high directivity, compact size, and reliable performance, differential microphone arrays (DMAs) have attracted great interest from both industry and academia as they have demonstrated great potential to be used in a wide range of applications for high-fidelity speech acquisition. Nevertheless, in many real-world applications, DMAs powered with fixed differential beamformers are often inadequate in suppressing interference, particularly in environments with multiple or moving sources. To address this issue, this work develops an adaptive convex combination (ACC)-based method, which combines multiple differential beamformers in an online manner for enhanced performance. While the major contribution is a new real-time processing algorithm that facilitates optimal linear combinations of different differential beamformers, making them adapted to dynamic environments, the presented method also provides valuable insights as how to combine different beamformers for online robust implementation. Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2024 | On the Design of Planar Differential Microphone Arrays with Specified Beamwidth or Sidelobe LevelabstractThis paper investigates the problem of designing differential beam-formers with planar microphone arrays to achieve not only the desired target directivity pattern but also control the beamwidth (BW) or sidelobe level (SLL). We first discuss the target directivity patterns and express the Dolph-Chebyshev polynomial based form of target directivity patterns into linear combination of cylindrical harmonics. We then address the problem of designing differential beamformers through beampattern approximation based on the Jacobi-Anger series expansion. Two methods are subsequently developed: the first one involves designing beamformers to achieve the target directivity pattern while minimizing SLL under a pre-specified value of BW and the second one aims to attain the target directivity pattern while minimizing the null-to-null BW under a pre-specified level of SLL. Simulations are carried out to validate the method and the results demonstrate the properties of the proposed method. Xueqin Luo, Jilu Jin, Gongping Huang, Yingke Zhao, Jingdong Chen, Jacob Benesty |
ICASSP | 6 |
| 2024 | A Steered Response Power Approach with Bilinear Prediction-Based Trade-Off Prewhitening for Speaker LocalizationabstractThis paper studies the problem of acoustic source localization in room environments. It presents an improved steered response power (SRP) approach with low-complexity and trade-off prewhitening. This method consists of two steps. In the first one, the linear predictor that is used to model the speech signals is formulated as a bilinear form, and a group of convex-constrained linear prediction sub-models with respect to dual sub-predictors are established to pre-filter microphone signals. The pre-filtered (prewhitened) microphone signals are subsequently used in SRP for speaker localization. Simulation results demonstrate the properties of the presented method: it is robust to reverberation and noise, and is computationally efficient thanks to the bilinear form. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 4 |
| 2024 | Directional Gain Based Noise Covariance Matrix Estimation for MVDR BeamformingabstractThis paper is devoted to the problem of noise covariance matrix (NCM) estimation. It proposes a time-frequency masking based approach. We first present an optimal mask function based on the mean-squared error criterion. To estimate this mask, we employ the recently developed directional gain method based on the knowledge of the signal incident angle. To demonstrate the effectiveness of the proposed NCM estimator, we integrate it into the minimum variance distortionless response (MVDR) beamformer. The speech enhancement results in noise-plus-interference environments show the advantages of the proposed method over two baseline beamforming algorithms. Fan Zhang 0001, Chao Pan 0001, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2024 | Differential Beamforming with Null Constraints for Spherical Microphone ArraysabstractDifferential microphone arrays (DMAs) can measure both the acoustic pressure field and the differential acoustic pressure fields, which gives them great advantages in a wide range of applications for acoustic and speech signal acquisition. The core component of DMAs is the so-called differential beamformer, the design of which typically involves taking into account the a priori knowledge about the array geometry and the desired directivity pattern that is related to the differential sound field to respond. This paper deals with the design of differential beamformers with spherical microphone arrays. It presents a novel design approach based on the null constraints formed from the desired directivity pattern. In comparison with the exiting methods, the proposed approach only requires the information of the zeros in the beampattern, which provides notable flexibility and convenience for spherical DMA design in practical applications. Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2024 | Nonlinear acoustic echo cancellation using low-complexity low-rank recursive least-squares algorithms
Vinal Patel, Sankha Subhra Bhattacharjee, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
Signal Process. | 5 |
| 2024 | On intrusive speech quality measures and a global SNR based metric
Chao Pan 0001, Jingdong Chen, Jacob Benesty |
Speech Commun. | 3 |
| 2024 | A Closed-Form DOA Estimator Using Spherical Microphone Arrays in the Presence of InterferenceabstractDirection-of-arrival (DOA) estimation is challenging in complex acoustic environments with background noise and interference. Utilizing spherical microphone arrays, closed-form estimators can be derived, which are attractive for practical applications due to their computational efficiency, eliminating the need for exhaustive extremum searching. However, current closed-form estimators are susceptible to interference. To address this issue, we propose an estimator that directly computes the DOA of the desired source using the covariance matrix of the observation signals. This approach effectively mitigates the impact of interference when the covariance matrix is accurately estimated. Simulation results demonstrate the superior performance of the proposed method compared to the subspace pseudo-intensity vector (SSPIV) and relative harmonic coefficients (RHC) methods. Yilong Lu, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 4 |
| 2024 | On the Design of Robust Differential Beamformers From the Beampattern Error PerspectiveabstractDifferential microphone arrays (DMAs), which enhance acoustic signals of interest by measuring both the acoustic pressure field and its spatial derivatives, find extensive use in various practical systems and acoustic products. A critical element of DMAs is the differential beamformer, traditionally designed to ensure that the designed beampattern closely matches the desired target directivity pattern. However, such beamformers may lack sufficient robustness in practice. To address the balance between robustness and beampattern accuracy, this letter proposes two types of beamformers: one prioritizes maximizing the white noise gain (WNG) while maintaining a specified mean-squared beampattern error (MSBE), and the other aims to minimize MSBE while adhering to a specified level of WNG. By transforming these design challenges into quadratic eigenvalue problems (QEPs), we derive explicit solutions for the proposed beamformers. Simulations are conducted to illustrate the performance characteristics of these beamformers. Jingli Xie, Junqing Zhang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 4 |
| 2024 | Smoothed Frame-Level SINR and Its Estimation for Sensor Selection in Distributed Acoustic Sensor NetworksabstractDistributed acoustic sensor network (DASN) refers to a sound acquisition system that consists of a collection of microphones randomly distributed across a wide acoustic area. Theory and methods for DASN are gaining increasing attention as the associated technologies can be used in a broad range of applications to solve challenging problems. However, unlike traditional microphone arrays or centralized systems, properly exploiting the redundancy among different channels in DASN is facing many challenges including but not limited to variations in pre-amplification gains, clocks, sensors' response, and signal-to-interference-plus-noise ratios (SINRs). Selecting appropriate sensors relevant to the task at hand is therefore crucial in DASN. In this work, we propose a speaker-dependent smoothed frame-level SINR estimation method for sensor selection in multi-speaker scenarios, specifically addressing source movement within DASN. Additionally, we devise an approach for similarity measurement to generate dynamic speaker embeddings resilient to variations in reference speech levels. Furthermore, we introduce a novel loss function that integrates classification and ordinal regression within a unified framework. Extensive simulations are performed and the results demonstrate the efficacy of the proposed method in accurately estimating smoothed frame-level SINR dynamically, yielding state-of-the-art performance. Shanzheng Guan, Mou Wang, Zhongxin Bai, Jianyu Wang 0007, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Design of Fully Steerable Differential Beamformers With Linear SuperarraysabstractLinear differential microphone arrays (LDMAs) are commonly integrated into thin and portable devices to achieve high-fidelity speech acquisition. Traditional LDMAs typically consist of only omnidirectional microphones, which impose limitations on their ability to produce steerable spatial responses due to constraints in array element directivity and linear array geometry. A recent solution to this limitation involves integrating both omnidirectional and bidirectional microphones in LDMA design, enabling the creation of steerable spatial responses. This paper extends the core idea of integrating omnidirectional and bidirectional microphones, and develops a more general and comprehensive theory and method for designing steerable LDMAs. It makes two main contributions. Firstly, it introduces a general approach to designing steerable LDMAs, in which any type of directional microphones can be used. Secondly, it gives the minimum number of omnidirectional and directional microphones required to achieve a specific order of steerable LDMA. Simulations validate the proposed method and illustrate how omnidirectional and directional sensors can be combined to form the desired LDMAs. Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Decomposition-Based Wiener Filter Using the Kronecker Product and Conjugate Gradient MethodabstractThe identification of long-length impulse responses represents a challenge in the context of many applications, like echo cancellation. Recently, the problem has been addressed in the framework of low-rank systems, using a decomposition of the impulse response based on the nearest Kronecker product and low-rank approximations. As a result, the original system identification problem that involves a long-length finite impulse response filter is reshaped as a combination of two (much) shorter filters, which leads to significant advantages. In this context, the benchmark Wiener filter can be formulated in terms of an iterative algorithm, where the estimates of the two component filters are sequently updated. However, matrix inversion operations are required within this algorithm. In this article, we develop a new version of the decomposition-based iterative Wiener filter, which relies on the conjugate gradient (CG) method and avoids matrix inversion. Simulations performed in the framework of echo cancellation indicate the good performance of the proposed solution, which outperforms the conventional Wiener filter (implemented using CG updates) and inherits the advantages of the decomposition-based approach. Cristian Lucian Stanciu, Jacob Benesty, Constantin Paleologu, Ruxandra-Liana Costea, Laura-Maria Dogariu, Silviu Ciochina |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | On Semi-Blind Source Separation-Based Approaches to Nonlinear Echo Cancellation Based on Bilinear Alternating OptimizationabstractAcoustic echo cancellation (AEC) is a crucial task in full duplex communications. As conventional linear filtering approaches are ineffective to deal with double-talk, various semi-blind source separation (SBSS)-based AEC algorithms are deceived, most of which are formulated and implemented in the frequency domain based on the multiplicative transfer function (MTF) model for computational efficiency. To avoid large latency and in order to deal with loudspeaker nonlinearities, the convolutive transfer function (CTF) model and odd power series expansion are leveraged, which are employed by numerous SBSS-based nonlinear AEC (SBSS-NAEC) algorithms. Conventional SBSS-NAEC methods estimate the series expansion coefficients and the CTF filter simultaneously making the number of free parameters to estimate large. Hence, the corresponding algorithms are computationally expensive and are difficult to optimize. In this work, we propose to decouple the series expansion coefficients and the CTF filters into a bilinear form and present a bilinear alternating optimization framework for estimating the model parameters. An alternating iterative projection (AIP) algorithm and an alternating element-wise iterative source steering (AEISS) algorithm are proposed. As the bilinear representation consists of less parameters compared to the conventional methods, the proposed algorithms not only improve the AEC performance but also reduce the computational complexity, which is validated by comprehensive simulations and experiments. Xianrui Wang, Yichen Yang 0010, Andreas Brendel, Tetsuya Ueda, Shoji Makino, Jacob Benesty, Walter Kellermann, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Interference-Controlled Maximum Noise Reduction Beamformer Based on Deep-Learned Interference ManifoldabstractBeamforming has been used in a wide range of applications to extract the signal of interest from microphone array observations, which consist of not only the signal of interest, but also noise, interference, and reverberation. The recently proposed interference-controlled maximum noise reduction (ICMR) beamformer provides a flexible way to control the specified amount of the interference attenuation and noise suppression; but it requires accurate estimation of the manifold vector of the interference sources, which is challenging to achieve in real-world applications. To address this issue, we introduce an interference-controlled maximum noise reduction network (ICMRNet) in this study, which is a deep neural network (DNN)-based method for manifold vector estimation. With densely connected modified conformer blocks and the end-to-end training strategy, the interference manifold is learned directly from the observation signals. This approach, akin to ICMR, adeptly adapts to time-varying interference and demonstrates superior convergence rate and extraction efficacy as compared to the linearly constrained minimum variance (LCMV)-based neural beamformers when appropriate attenuation factors are selected. Moreover, via learning-based extraction, ICMRNet effectively suppresses reverberation components within the target signal. Comparative analysis against baseline methods validates the efficacy of the proposed method. Yichen Yang 0010, Ningning Pan, Wen Zhang 0002, Chao Pan 0001, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | A Frequency-Domain Recursive Least-Squares Adaptive Filtering Algorithm Based On A Kronecker Product DecompositionabstractThis paper proposes a frequency-domain recursive least-squares (RLS) adaptive filtering algorithm for identifying time-varying acoustic systems in noisy environments. The Kronecker product (KP) is employed to decompose the model filter of the acoustic channel impulse response into two sets of short sub-filters, based on which a generalized frequency-domain signal model and the associated cost function are established. A KP based RLS algorithm is subsequently deduced. In comparison with the conventional frequency-domain RLS adaptive filter, the presented algorithm is not only computationally more efficient, but also has a faster convergence rate for the identification of acoustic systems regardless of whether the excitation is a white sequence or a speech signal. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 3 |
| 2023 | Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech DereverberationabstractDereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those methods generally assume that there is only a single speaker in the acoustic environment and, consequently, they suffer from significant performance degradation if multiple speakers participate in the conversation. How to deal with reverberation in multiple-speaker scenarios is still a challenging problem, which is studied in this work. We present a switching multichannel linear prediction filtering method, which designs multiple linear filters with each tracking one speaker. When some speaker is active, the corresponding filter and the weighted cross-correlation matrix are updated while the other filters are kept unchanged. To further improve the performance and reduce complexity, we apply the Kronecker product to decompose every linear prediction filter into a Kronecker product of two shorter filters: one is time-invariant and the other is time-varying. The former is estimated with a batch method (using only a few seconds of speech signal when the corresponding speaker starts to talk in the entire conversation) while a recursive least-squares algorithm is derived for identifying the time-varying set of Kronecker filters. Gongping Huang, Jacob Benesty, Israel Cohen, Emil Winebrand, Jingdong Chen, Walter Kellermann |
ICASSP | 2 |
| 2023 | On Multiple-Input/Binaural-Output Antiphasic Speaker Signal ExtractionabstractThis paper studies the problem of target speaker signal exaction and antiphasic rendering with an array of microphones in the scenarios where there are two active speakers. Based on the important findings achieved in the psychoacoustic field as well as our recent works on single-channel speech enhancement, we present a rendering based approach in which a temporal convolutional network (TCN) is trained to take the multiple signals observed by the microphone array as its inputs and generate two output (binaural) signals. The TCN is trained in such a way that, when binaural output signals are listened by the listener with headsets, the speech signal from the desired speaker is perceived on one side of and close to the listener’s head, while the competing speech signal is perceived on the opposite side and also away from the listener’s head. Benefited from rendering and the signal-to-interference ratio (SIR) improvement, this antiphasic binaural presentation enables the listener to better focus on the target speaker’s signal while ignoring the impact of the competing speech. The modified rhyme tests (MRTs) are performed to validate the superiority of the proposed method. Xianrui Wang, Ningning Pan, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2023 | A binaural heterophasic adaptive beamformer and its deep learning assisted implementation
Jilu Jin, Ningning Pan, Jingdong Chen, Jacob Benesty, Yiqian Yang |
Pattern Recognit. Lett. | 4 |
| 2023 | Recursive least-squares algorithm based on a third-order tensor decomposition for low-rank system identification
Constantin Paleologu, Jacob Benesty, Cristian Lucian Stanciu, Jesper Rindom Jensen, Mads Græsbøll Christensen, Silviu Ciochina |
Signal Process. | 2 |
| 2023 | Dimensionality Reduction of Room Acoustic Impulse Responses and Applications to System IdentificationabstractA room Acoustic Impulse Response (RAIR), which represents the sound propagation channel via direct and reflection paths from a source position to a microphone, plays a leading role in a broad range of acoustic signal processing applications, e.g., echo cancellation. In practical acoustic environments, it is not uncommon that an RAIR may consist of hundreds or even thousands of coefficients, making it challenging to identify and handle. This paper investigates the RAIR dimensionality reduction problem inspired from the concepts of dynamic mode decomposition. The objective is to find effective lower-dimensional representations of RAIRs, which are easier and more robust to identify and equalize. There are two main contributions of this work. First, we present an RAIR dimensionality reduction method. Second, we show how to apply this technique to the problem of acoustic system identification. Simulation results demonstrate that the proposed method is able to improve significantly the performance of acoustic system identification. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 2 |
| 2023 | Differential Beamforming From a Geometric PerspectiveabstractDifferential microphone arrays (DMAs) have demonstrated a great potential for solving the high-fidelity sound acquisition problem in a wide range of applications as they possess many good properties such as frequency-independent beampatterns with high directivity. A significant number of efforts have been devoted to the design of DMAs and the associated beamformers. As a result, many different types of DMAs and differential beamforming methods have been developed over the last few decades, some of which have been successfully deployed in real systems and commercial products. However, given an application, how to design a DMA to achieve optimal performances is still an open issue. This work studies the problem of designing linear DMAs (LDMAs) from a geometric perspective. Based on the fundamental observation that most practical and interesting DMA beampatterns have nulls in some directions, we define a criterion based on the orthogonality between the beamforming filter and the steering vector in the nulls' directions. We then derive a family of differential beamformers by optimizing the defined criterion, some of which are well known but derived from a different perspective, while others are new. Simulations and experiments are carried out, and the results validate the proposed method and developed differential beamformers. Jilu Jin, Jacob Benesty, Jingdong Chen, Gongping Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Design of Maximum Directivity Beamformers With Linear Acoustic Vector Sensor ArraysabstractThis paper studies the design of maximum directivity factor (MDF) beamformers based on uniform linear arrays (ULAs) consisting of acoustic vector sensors (AVSs). We first derive the main lobe constraints, which ensure that the beamformer's beampattern achieves a maximum in the look direction, and prove that any beamformer that satisfies the proposed constraints can be written as the sum of two orthogonal beamformers: the maximum white noise gain (MWNG) beamformer and a reduced-rank beamformer. Then, we derive the MDF beamformer by maximizing the directivity factor (DF) under the deduced constraints. We also derive a robust version of the MDF beamformer, which can keep the WNG above a pre-specified level. Compared to the conventional MDF beamformer based on ULAs with omnidirectional microphones, the designed MDF beamformer with uniform linear AVS arrays (ULAVSAs) can steer the beampattern to any look direction in the 3-dimensional space and achieves a higher directivity. The proposed MDF beamformer also outperforms the two-step MDF beamformer with ULAVSAs since it maximizes the DF. The proposed methods are validated through simulations as well as real experiments. Xueqin Luo, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty, Wen Zhang 0002, Mengyao Zhu 0003, Chunjian Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Design of 2D and 3D Differential Microphone Arrays With a Multistage FrameworkabstractDifferential microphone arrays (DMAs) have demonstrated a great potential for high-fidelity acoustic and speech signal acquisition in a wide range of applications since such arrays are able to achieve frequency-invariant beampatterns with high directivity. Consequently, a great number of efforts have been devoted to the design of DMAs and the associated beamformers in the literature. However, most of the methods only work for arrays with particular topologies, e.g., linear, circular, concentric circular, and spherical ones. How to design general two-dimensional (2D) and three-dimensional (3D) DMAs that can measure the desired differential sound field and form the desired spatial response in the 3D space remains an unsolved problem. This paper investigates this problem and presents a multistage design approach. The major contributions of this work are as follows. First, we reexamine the differentials of the acoustic pressure field in the 3D space and derive the general expression of the directivity patterns resulting from the spatial differential operation, which serves as the foundation for differential beamforming with 2D or 3D microphone arrays. Second, we present a multistage approach to the design of 2D and 3D DMAs, and deduce the relationship between the global beamformer and the beamformers at different stages as well as the relationship between their beampatterns. Third, several algorithms are presented for the design of differential as well as robust differential beamformers in this multistage framework. Simulation results validate the proposed approach and justifies its properties. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | LMS and NLMS Algorithms for the Identification of Impulse Responses with Intrinsic Symmetric or Antisymmetric PropertiesabstractIn applications involving system identification problems, some characteristics of the impulse response of the system to be identified are usually exploited to design adaptive algorithms with improved performance. In this context, this paper focuses on the identification of systems that own intrinsic symmetric or antisymmetric properties, which can be further formulated by using a combination of bilinear forms. Based on such an approach, the least-mean-square (LMS) and normalized LMS (NLMS) algorithms with symmetric/antisymmetric properties (termed here LMS-SAS and NLMS-SAS) are proposed. Simulation results are shown confirming the improved convergence speed achieved by the proposed algorithms as compared to the conventional LMS and NLMS counterparts for different operating scenarios. Jacob Benesty, Constantin Paleologu, Silviu Ciochina, Eduardo Vinicius Kuhn, Khaled Jamal Bakri, Rui Seara |
ICASSP | 1 |
| 2022 | DNN Based Multiframe Single-Channel Noise Reduction FiltersabstractWhile multiframe noise reduction filters, e.g., the multiframe Wiener and minimum variance distortionless response (MVDR) ones, have demonstrated great potential to improve both the subband and full-band signal-to-noise ratios (SNRs) by exploiting explicitly the interframe speech correlation, the implementation of such filters requires the knowledge of the interframe correlation coefficients for every subband, which are challenging to estimate in practice. In this work, we present a deep neural network (DNN) based method to estimate the interframe correlation coefficients and the estimated coefficients are subsequently fed into multiframe filters to achieve noise reduction. Unlike existing DNN based methods, which outputs the enhanced speech directly, the presented method combines deep learning and traditional methods, which gives more flexibility to optimize or tune noise reduction performance. Experimental results are presented to justify the properties of the presented methods. Ningning Pan, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2022 | Study of the Null Directions on The Performance of Differential BeamformersabstractNull directions are important parameters for differential beamformers, which play an important role on the beamforming performance. In this paper, we investigate the performance of differential beamformers as a function of the null directions. We first derive the directivity factor (DF) as an explicit function of null and show that the DF decreases to 0 if any null approaches to the desired look direction. We then validate the theoretical analysis through simulations using the beampattern, DF and signal-to-interference gain as the performance measures. The results show that: 1) the performance of a differential beamformer degrades significantly if there is any null close to the desired look direction; 2) with a fixed null direction, increasing the order of the differential beamformer can help improve performance. Xuehan Wang, Israel Cohen, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2022 | Multistage approach for steerable differential beamforming with rectangular arrays
Gal Itzhak, Jacob Benesty, Israel Cohen |
Speech Commun. | 2 |
| 2022 | Data-Reuse Recursive Least-Squares AlgorithmsabstractThere are different strategies to improve the overall performance of the recursive least-squares (RLS) adaptive filter. In this letter, we focus on the data-reuse approach, aiming to improve the convergence rate/tracking of the algorithm by reusing the same set of data (i.e., the input and reference signals) several times. First, we present a computationally efficient data-reuse RLS algorithm, which is the result of a low complexity implementation of the data-reuse process. Moreover, we extend the idea to the fast RLS algorithm. Simulations performed in the context of echo cancellation support the performance gain. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
IEEE Signal Process. Lett. | 2 |
| 2022 | Identification of Room Acoustic Impulse Responses via Kronecker Product DecompositionsabstractThe identification of room acoustic impulse responses represents a challenging problem in the framework of many important applications related to the acoustic environment, like echo cancellation, noise reduction, and microphone arrays, among others. In this context, the main issues are related to the long length of such impulse responses and their time-variant nature. These raise significant difficulties in terms of the convergence rate, computational complexity, and accuracy of the solution. Recently, a decomposition-based approach was developed for the identification of low-rank systems, which can also be applied (to some extent) for the identification of acoustic impulse responses. This approach exploits the nearest Kronecker product decomposition of the impulse response and solves a high-dimension system identification problem using a combination of low-dimension solutions (provided by shorter filters), thus gaining in terms of both performance and complexity. Nevertheless, it does not consider the intrinsic nature of the room acoustic impulse responses, which contain specific components (e.g., early reflections and late reverberation) that can be very different in nature. In this paper, we propose an improved decomposition-based method (via the Kronecker product) that takes into account these specific components and processes them separately, in order to better exploit their important low-rank features. Following this approach, an iterative Wiener filter is firstly developed, followed by a recursive least-squares (RLS) algorithm designed in the same framework. Both solutions outperform the conventional benchmarks, i.e., the conventional Wiener filter and the RLS algorithm, respectively. Moreover, they achieve superior performances as compared to the recently developed versions based on the nearest Kronecker product decomposition, also owning lower computational complexities than their previous counterparts. Simulations are performed in the framework of acoustic echo cancellation and the obtained results support the performance features of the proposed algorithms. Laura-Maria Dogariu, Jacob Benesty, Constantin Paleologu, Silviu Ciochina |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Fundamental Approaches to Robust Differential Beamforming With High Directivity FactorsabstractDifferential beamforming, which measures the spatial derivatives of the acoustic pressure field, can be used in a wide range of small devices that require high-fidelity sound and speech acquisition as it can achieve frequency-invariant spatial responses with high directivity factors (DFs). Since a differential process is inherently sensitive to sensors' self noise and other array imperfections, the most challenging problem in the design of any differential beamformer is how to achieve the maximum possible DF while maintaining a proper level of robustness for practical usage. While significant efforts have been made on this topic, the problem remains unsolved and further study is indispensable. This paper is devoted to dealing with this challenging problem. It presents a study on theory and methods to achieve the optimal and fundamental compromise between the white noise gain (WNG), which quantifies how robust is the beamformer, and the DF in differential beamforming. The major contributions of this work are as follows. 1) We show and prove that any null constrained fixed beamformer can be decomposed as the sum of two orthogonal filters, i.e., the maximum WNG (MWNG) beamformer and a reduced-rank one. Based on this decomposition, we develop three kinds of differential beamformers from the WNG perspective, which can achieve a flexible and optimal compromise between DF and WNG. 2) We show that a transformed null constrained beamformer can also be decomposed as the sum of two orthogonal filters, i.e., the transformed maximum DF (MDF) beamformer and another reduced-rank one. Based on this decomposition, we also develop three kinds of differential beamformers, which can obtain the desired level of DF while using the rest of the degrees of freedom to maximize the WNG. Simulations are performed to validate the theoretical analysis and developed differential beamformers. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Kronecker Product Multichannel Linear Filtering for Adaptive Weighted Prediction Error-Based Speech DereverberationabstractReverberation, whichis caused by late reflections, impairs not only speech quality but also intelligibility. Consequently, dereverberation, a process to mitigate the impact of reverberation, has attracted significant research interests. Numerous approaches have been developed in the literature, among which the weighted-prediction-error (WPE) one has demonstrated promising potential for reducing or eliminating reverberation. The WPE method has been well studied and several variants have been developed. The adaptive one, called adaptive WPE (AWPE) method, has been widely investigated for use in real applications as it can deal with reverberation in time-varying acoustic environments. However, the computational complexity of AWPE is high, which may be a problem for its implementation in real-time systems. This paper presents some new insights into AWPE-based speech dereverberation by introducing the concepts of Kronecker product and partially time-varying filtering. It then develops two algorithms for dereverberation with lower complexity than AWPE. The significant contributions of this work are as follows. First, we propose a Kronecker product filtering framework for speech dereverberation, where the linear prediction filter is formulated as the Kronecker product of two sets of shorter filters. Second, we propose a partially time-varying Kronecker product filter for dereverberation. Instead of estimating the entire linear prediction filter as in the conventional method, the proposed one only needs to update part of the filter. The proposed approaches can significantly reduce the computational complexity without sacrificing dereverberation performance as compared to AWPE. Simulation results validate the theoretical analysis and justify the advantages of the new methods. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | On Differential Beamforming With Nonuniform Linear Microphone ArraysabstractWhile differential beamforming with uniform linear arrays (ULAs) has been widely studied, there is little work so far regarding the design of differential beamformers with nonuniform linear arrays (NULAs). This paper attempts to shed some light on the principles of differential beamforming with NULAs. We define spatial difference operators with NULAs, where any order of the spatial difference of the observation signals can be represented as the product of a nonuniform spatial difference operator matrix and the observation vector. Consequently, the design of differential beamformers is performed in two stages. In the first one, a nonuniform spatial difference operator matrix is applied to the array observations, thereby yielding differential signals. In the second stage, beamformers are designed and applied to the obtained differential signals to optimize the array performance. Based on the defined spatial difference operators, we derive from some performance metrics a family of differential beamformers with NULAs, which include the maximum directivity factor (DF), the maximum white noise gain (WNG), and the maximum front-to-back ratio (FBR) differential beamformers. To compromise between the DF and array robustness, we also derive the parameterized maximum DF and parameterized maximum FBR differential beamformers. The null-constraint maximum DF and WNG differential beamformers are also developed so that some nulls can be placed in specified directions for interference suppression. Simulation results validate the theoretical analysis and justify the properties of the proposed methods. Jilu Jin, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Microphone Array Beamforming With High Flexible Interference Attenuation and Noise ReductionabstractThis paper studies the problem of microphone array beamforming to enhance a speech signal of interest in adverse acoustic environments, where interference and additive background noise coexist. The problem is formulated as one of convex optimization whose solution under a specified level of interference attenuation leads to an interference controlled maximum noise reduction (ICMR) beamformer, which can be expressed as a linear combination of two MVDR beamformers: one attempts to extract the desired source signal while the other attempts to extract the interference. The combination coefficients are functions of the array manifold vectors, noise coherence matrix, and the specified interference attenuation factor. By tuning the interference attenuation factor, the ICMR beamformer can be implemented to achieve aggressive interference attenuation or even eliminate interference completely; but this may lead to less additive noise suppression or even noise amplification. To control the maximum sacrifice in gain (SG) of the signal-to-noise ratio (SNR) that is acceptable for additive reduction, a variant of ICMR is derived, which is named as the ICMR-SG beamformer. Simulations are performed and the results show that the ICMR beamformer is able to control the amount of interference attenuation. In comparison, ICMR-SG controls the maximum SG of SNR while achieving the optimal possible level of interference attenuation. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Planar Array Geometry Optimization for Region Sound AcquisitionabstractMicrophone arrays have been used in wide range of applications for sound acquisition and signal enhancement, the performance of which depends not only on the processing algorithms but also on the array geometry. A large number of efforts have been devoted to the development of beamforming and signal enhancement algorithms for processing microphone array signals in the literature. Relatively, few efforts have been made to investigate the problem of array geometry optimization. This paper studies the problem of geometry optimization for planar arrays and it develops a genetic optimization algorithm that can optimize the positions of the sensors, thereby maximizing the directivity factor (DF) with a constrained level of white noise gain (WNG) given the number of microphones, the region in which they should be placed, and the interested range of steering. Simulation results show that the optimized array geometry outperforms the uniform linear, the uniform circular and the rectangular grid geometries in terms of DF with the same number of sensors and the same constraint on the minimum level of WNG. Xi Chen 0128, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2021 | Robust Recursive Least M-Estimate Adaptive Filter for the Identification of Low-Rank Acoustic SystemsabstractTo identify acoustic systems (which are low-rank in nature) in non-Gaussian and Gaussian noise, a robust recursive least M-estimate adaptive filtering algorithm is developed in this paper by applying the nearest Kronecker product to decompose the acoustic impulse response. Two M-estimators, i.e., the Cauchy and Welsch estimators, are employed to define the cost function of the adaptive filter, leading to a class of numerically stable adaptive filtering algorithms, which are robust to non-Gaussian noise. The effectiveness of the developed algorithm is validated in acoustic environments with both Gaussian and non-Gaussian noise. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 3 |
| 2021 | Combined Differential Beamforming With Uniform Linear Microphone ArraysabstractWhile differential beamformers have been widely used in voice communication and human-machine speech interface systems to enhance speech signals of interest, how to design such beamformers that on the one hand can achieve the highest possible directivity factor (DF) and on the other hand are able to obtain a certain level of white noise gain (WNG), so that they are robust enough to sensors’ self noise and array imperfections is still a challenging issue. This paper studies the problem of robust differential beamforming with small-size arrays to achieve a high DF. It presents a method for the design of differential beamformers with uniform linear arrays. We first generate differential pressure signals by applying the recently developed forward spatial difference operator to the outputs of the array with pressure sensors. The pressure microphone observation signals and the differential pressure signals are then put together, and a combined beamformer is subsequently designed, which consists of two subbeamformers, one operates on the pressure microphone observations and the other on the differential pressure signals. A new class of combined differential beamformers are introduced, which can achieve different levels of compromises between DF and WNG using an adjustable parameter. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
ICASSP | 3 |
| 2021 | Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone ArraysabstractDifferential beamformers with concentric circular microphone arrays (CCMAs) are desirable for use in various applications since they can form frequency-invariant spatial responses, have better beam steering flexibility than linear arrays, and suffer less with beampattern irregularity and white noise amplification than circular microphone arrays (CMAs). The methods developed previously for differential beamforming with CCMAs are based on the series expansion. Such methods need to know the analytic form of the target beam-pattern, which may not be accessible in practice. Furthermore, expansion error may lead to erroneous solution, which can cause noise amplification instead of reduction. In this paper, we extend our recently developed beamforming method for CMAs to the design of differential beamformers with CCMAs, which takes advantage of the symmetric null constraints from the beampattern. Simulations are performed to justify the properties of the proposed approach. Xuehan Wang, Gongping Huang, Israel Cohen, Jacob Benesty, Jingdong Chen |
ICASSP | 4 |
| 2021 | A Simplified Wiener Beamformer Based on Covariance Matrix ModellingabstractThis paper is devoted to the problem of adaptive beamforming with small-spaced microphone arrays. In this context, the Wiener filter is an optimal beamformer in the mean-squared error (MSE) sense. However, it requires good estimates of the covariance matrices of the speech signal of interest and noise, which are difficult to achieve in time-varying and reverberant acoustic environments. To deal with this problem, we propose a general method by parametric modeling the covariance matrices of speech and noise, which leads to a simplified Wiener beamformer. This beamformer has only one time-varying parameter to estimate, which is much easier to achieve as compared to the estimation of covariance matrices. As an example, we adopt the parametric model used in the superdirective beamformer, which models the covariance matrices as a combination of the pseudo-coherence matrices of a point source and diffuse noise. Simulation results show that the developed beamformer outperforms the traditional Wiener beamformer in terms of both noise and reverberation suppression. Fan Zhang 0001, Chao Pan 0001, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2021 | On the Design of Square Differential Microphone Arrays with a Multistage StructureabstractThis paper studies the problem of designing square differential microphone arrays (SDMAs). It presents a multistage approach, which first divides an SDMA composed of M2microphones into (M − 1)2subarrays with each subarray being a 2 × 2 square array formed by four adjacent microphones. Then, differential beamforming is performed with each subarray in the first-stage. The first-stage differential beamformers’ outputs are subsequently used as the inputs of the second stage to form (M − 2)2subarrays and a second-stage differential beamforming is then performed. Continuing this process till the (M −1)th stage, we obtain the final output of the SDMA. The SDMA designed in such a multistage structure has two important properties. First, the global weighting matrix is equal to the two dimensional convolution of weighting matrices from the first stage to the last one. Second, the global beampattern is equal to the product of beampatterns from all stages. Consequently, we can combine different kinds of beamformers in different stages and have better control of the performance metrics. Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
ICASSP | 3 |
| 2021 | Adaptive line enhancer for nonstationary harmonic noise reduction
Aviva Atkins, Israel Cohen, Jacob Benesty |
Comput. Speech Lang. | 3 |
| 2021 | A New Method to Design Steerable First-Order Differential BeamformersabstractFirst-order differential microphone arrays (FODMAs), which combine a small-spacing uniform linear array and a first-order differential beamformer, have been used in a wide range of applications for sound and speech signal acquisition. However, traditional FODMAs are not steerable and their main lobe can only be at the endfire directions. To circumvent this problem, we propose in this letter a new method to design steerable FODMAs. We first divide the target beampattern into a sum of two sub-beampatterns, i.e., cardioid and dipole, where the summation is controlled by the steering angle. We then design two sub-beamformers, one is similar to the traditional approach and is used to achieve the cardioid sub-beampattern, while the other is designed to filter the squared observation signals and is used to approximate the dipole sub-beampattern. The overall beampattern resembles the target beampattern for any steering angle. Simulations and experiments are performed to justify the effectiveness of the developed method. Xin Leng, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2021 | A Single-Input/Binaural-Output Antiphasic Speech Enhancement Method for Speech Intelligibility ImprovementabstractImproving intelligibility of a speech signal of interest from its observations (with a single microphone) corrupted by additive noise has long been a challenging problem. Motivated by important findings achieved in the psychoacoustic field, we propose in this work a deep learning based method to render the noise and desired speech in the perceptual space such that the perception of the desired speech is least affected by the noise. Specifically, we adopt the temporal convolutional network (TCN) based structure to map the single-channel noisy observations into two binaural signals, one for the left ear and the other for the right ear. The TCN is trained in such a way that the desired speech and noise will be perceived to be in opposite directions when the listener listens to the binaural signals. This antiphasic binaural presentation enables the listener to better distinguish the desired speech from the annoying noise for improved speech intelligibility. The modified rhyme test is performed for evaluation and the results justify the superiority of the proposed method for speech intelligibility improvement. Ningning Pan, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 4 |
| 2021 | Time Difference of Arrival Estimation Based on a Kronecker Product DecompositionabstractTime difference of arrival (TDOA) estimation, which often serves as the fundamental step for a source localization or a beamforming system, has a significant practical importance in a wide spectrum of applications. To deal with reverberation, the TDOA estimation problem is often transformed into one of identifying the relative acoustic impulse responses. This letter presents a method to efficiently identify the relative acoustic impulse response between two microphones for TDOA estimation based on the so-called Kronecker product decomposition. By decomposing the relative impulse response into a series of Kronecker products of shorter filters, the original channel identification problem with a long impulse response is converted into one of identifying a number of short filters. Since the TDOA information is embedded only in the direct path of the relative impulse response, the dimension of the Kronecker product decomposition can be very small and, as a result, the developed algorithm is expected to work well in real environments with a small number of data snapshots. Xianrui Wang, Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
IEEE Signal Process. Lett. | 3 |
| 2021 | Robust Dereverberation With Kronecker Product Based Multichannel Linear PredictionabstractReverberation impairs not only the speech quality, but also intelligibility. The weighted-prediction-error (WPE) method, which estimates the late reverberation component based on a multichannel linear predictor, is by far one of the most effective algorithms for dereverberation. Generally, the WPE prediction filter in every short-time-Fourier-transform (STFT) subband has to be long enough to estimate accurately the late reverberation component. As a consequence, WPE is computationally expensive, which makes it difficult to implement into real-time embedded or edge computing devices. Moreover, WPE is sensitive to additive noise and its performance may suffer from dramatic degradation even in environments where the signal-to-noise ratio (SNR) is high. To address these drawbacks, this letter proposes to decompose the multichannel linear prediction filter as a Kronecker product of a temporal (interframe) prediction filter and a spatial filter. An iterative algorithm is then developed to optimize the two filters. In comparison with the original WPE algorithm, the presented method not only exhibits better performance in terms of dereverberation and robustness to additive noise, as there are fewer parameters to estimate for a given number of observation signal samples, but is also computationally more efficient, since the dimensions of the covariance matrices after Kronecker product decomposition are smaller. Wenxing Yang, Gongping Huang, Jingdong Chen, Jacob Benesty, Israel Cohen, Walter Kellermann |
IEEE Signal Process. Lett. | 4 |
| 2021 | On a Particular Family of Differential Beamformers With Cardioid-Like and No-Null PatternsabstractDifferential microphone arrays (DMAs), which are responsive to the differential acoustic pressure fields, have been used in a wide range of applications related to audio and speech. The core part of a DMA is the so-called differential beamformer, which is generally designed by placing a number of nulls in its beampattern to attenuate noise from some directions. But the presence of these nulls may cause some great issues, e.g., leading to suboptimal performance if the interference/noise is incident from directions other than the nulls' directions, and making the beamformer less robust to sensors' self noise and array imperfections. To overcome these problems, this letter is devoted to the design of differential beamformers with no nulls in its beampattern. A design method and its multistage implementation are presented and analyzed. An improved solution is then developed, which is able to form frequency-invariant beampatterns with no nulls in the frequency range of speech signals. Simulations are provided to illustrate the properties of the developed methods. Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE Signal Process. Lett. | 2 |
| 2021 | On the Robustness of the Superdirective BeamformerabstractIn microphone array beamforming, a high directional gain is always desired for acoustic noise and reverberation suppression; as a result, the superdirective beamformer has been of great interest in many applications. However, this beamformer is well known to be very sensitive to array imperfections. While much effort has been made to improve its robustness, it is still a major problem. This paper is essentially devoted to the study of the robustness of the superdirective beamformer and derivation of better ways to deal with this important issue. We first prove that any distortionless fixed beamformer can be written as the sum of two orthogonal beamformers, i.e., the sum of the classical delay-and-sum (DS) beamformer and a reduced-rank beamformer. Based on this property, different kinds of robust superdirective beamformers are then developed. We also show that the robust design problem can be transformed into a quadratic eigenvalue problem (QEP), which leads to a solution that achieves the maximum possible directivity factor (DF) while meets the white noise gain (WNG) constraint over a frequency band of interest. Xi Chen 0128, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | On the Design of Differential Kronecker Product BeamformersabstractIn this paper, we present a generalized approach for differential microphone array (DMA) beamforming in the short-time Fourier transform (STFT) domain. We propose a multistage beamforming approach, which considers a Kronecker product (KP) decomposition of the global beamformer into two independent sub-beamformers. We derive differential KP beamformers according to different criteria and analyze their performances, which are tuned by three design parameters. These parameters allow a high beamforming design flexibility; in particular, non-differential or non-KP beamformers may be obtained as special cases. Depending on the selection of parameters, we demonstrate a preferable performance with the new approach with respect to the white noise gain and directivity factor measures. In addition, we consider the task of speech enhancement. We show that differential KP beamformers perform better than non-differential and non-KP beamformers in terms of the quality and intelligibility of their respective time-domain enhanced signals, particularly in moderately reverberant environments. Gal Itzhak, Jacob Benesty, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Steering Study of Linear Differential Microphone ArraysabstractDifferential microphone arrays (DMAs) can achieve high directivity and frequency-invariant spatial response with small apertures; they also have a great potential to be used in a wide spectrum of applications for high-fidelity sound acquisition. Although many efforts have been made to address the design of linear DMAs (LDMAs), most developed methods so far only work for the situation where the source of interest is incident from the endfire direction. This paper studies the steering problem of differential beamformers with linear microphone arrays. We present new insights into beam steering of LDMAs and propose a series of steerable differential beamformers. The major contributions of this paper are as follows. 1) A series of ideal functions are defined to describe the ideal, target beampatterns of LDMAs. 2) We prove that first-order differential beamformers with linear microphone arrays are not steerable and their mainlobes can only be at the endfire directions. 3) We deduce the fundamental conditions for designing steerable differential beamformers with LDMAs. 4) We develop a method to design steerable beamformers with LDMAs using null constraints. Simulations and experiments validate the properties of the developed method. Jilu Jin, Gongping Huang, Xuehan Wang, Jingdong Chen, Jacob Benesty, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Beamforming with Cube Microphone Arrays Via Kronecker Product DecompositionsabstractMicrophone arrays combined with beamforming have been widely used to solve many important acoustic problems in a wide range of applications. Much effort has been devoted in the literature to microphone array beamforming, among which the Kronecker product beamforming method developed recently has demonstrated some interesting properties. Generally, this method decomposes the global beamforming filter into a Kronecker product of a number of sub-beamforming filters, each of which corresponds to a virtual subarray and can be designed individually. This decomposition not only reduces significantly the number of beamforming coefficients, but also can be explored to improve the robustness and flexibility of beamforming. This paper extends Kronecker product beamforming from two-dimensional arrays into three-dimensional cube arrays. We consider two decompositions, i.e., fully and partially separable ones. The former decomposes the entire array into three linear subarrays while the latter decomposes the entire array into a linear subarray and a planar one. Then, for each case, we derive the Kronecker product maximum white noise gain beamformer, the Kronecker product approximate maximum directivity factor (DF) beamformer, the Kronecker product null-steering beamformer, and the Kronecker product iterative maximum DF beamformer. Simulation results demonstrate the properties and advantages of the proposed beamformers. Xuehan Wang, Jacob Benesty, Jingdong Chen, Gongping Huang, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | A New Class of Differential BeamformersabstractDifferential microphone arrays (DMAs) have been used in a wide range of applications for high-fidelity acoustic signal acquisition and enhancement. In the design of differential beamformers, three of the widely used measures are the directivity factor (DF), the front-to-back ratio (FBR), and the white noise gain (WNG). The former two have been used to obtain optimal differential beamformers, e.g., the hypercardioid and supercardioid, and the third one is generally used to analyze and control the robustness of the beamformer with respect to array imperfections due to sensors' self noise, mismatch among sensors, and sensors' placement errors. In this paper, we present a new measure called directivity factor and front-to-back ratio (DFBR), which is a generalization of DF and FBR. With this new measure, three different kinds of beamformers are derived. The first one is the maximum DFBR beamformer, which is deduced by maximizing DFBR with a joint diagonalization method. The second one is the ψ-cardioid beamformer, which is the maximum DFBR beamformer corresponding to a distortionless constraint. The last one is the reduced-rank differential beamformer, which is obtained by properly choosing the dimension of the signal subspace and maximizing WNG subject to the distortionless constraint. The developed beamformers have many interesting properties, which are justified by both simulations and experiments. Wenxing Yang, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Differential Beamforming From the Beampattern Factorization PerspectiveabstractDifferential beamformers have demonstrated a great potential in forming frequency-invariant beampatterns and achieving high directivity factors. Most conventional approaches design differential beamformers in such a way that their beampatterns resemble a desired or target beampattern. In this paper, we show how to design differential beamformers by simply taking advantage of the fact that the beampattern is actually a particular form of an exponential polynomial. Thanks to this quite obvious formulation, a target beampattern is not really needed while the zeros of the exponential polynomial and/or its factorization are fully exploited. The advantage of this factorization is twofold. First, it gives the relation between the beamformer and the roots of the polynomial, so the former can be directly determined from the latter, which seems natural and convenient. Second, based on this factorization, we propose a new formulation of the beamforming filter, which decomposes the filter into shorter ones with the Kronecker product. This formulation is very general, and many well-known beamformers such as the differential and delay-and-sum (DS) ones can be derived from it. Furthermore, the new formulation allows one to combine different kinds of beamformers together, which gives a great flexibility in forming different beampatterns and achieving a compromise among the directivity factor (DF), white noise gain (WNG), and frequency invariance. Jacob Benesty, Jingdong Chen, Gongping Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | On the Design of 3D Steerable Beamformers With Uniform Concentric Circular Microphone ArraysabstractCircular microphone arrays (CMAs) and concentric CMAs (CCMAs) have been used in a wide range of applications such as smartspeakers and teleconferencing systems because of their flexible steering ability. Although many efforts have been devoted to beamforming with CCMAs, most existing methods consider only the 2-dimensional (2D) case and assume that the sound sources of interest are in the same plane as the sensor array (generally the horizontal plane), which often does not hold true in practical applications. This paper deals with the problem of beamforming with uniform CCMAs (UCCMAs) in the 3-dimensional (3D) space to control the steering of the spatial response and meanwhile form frequency-invariant beampatterns for processing broadband acoustic and speech signals. The major contributions of this work are summarized as follows: 1) it presents an analysis based on the spherical harmonics decomposition about the Nth-order optimal and steerable directivity patterns; 2) a beamforming method is developed in which the beamformer's coefficients are identified by solving a linear system of equations formed by approximating the Nth-order optimal target beampattern with the beamformer's beampattern while the resulting beampattern can be steered flexibly in the 3D space; 3) the sufficient and necessary condition on the array geometry and sensors' placement are given to ensure that the beamformer exists and is unique; and 4) the analytical forms of the directivity factor (DF) and white noise gain (WNG) of the resulting beamformer is given and discussion is presented on what conditions irregularities (deep nulls) in WNG and DF may occur. Simulations are provided to illustrate the property of the developed beamforming methods. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Bilinear Models for Machine Learning
Tayssir Doghri, Leszek Szczecinski, Jacob Benesty, Amar Mitiche |
ICANN (1) | 3 |
| 2020 | Robust Frequency-Domain Recursive Least M-Estimate Adaptive Filter For Acoustic System IdentificationabstractTo identify acoustic systems in non-Gaussian and Gaussian noises, a robust frequency-domain recursive least M-estimate (FRLM) adaptive filtering algorithm is proposed. The cost function of the adaptive filter is defined by using a robust time-domain M-estimator, while its update equation is derived from the normal equation in the frequency domain. As compared to the frequency-domain recursive least-squares adaptive filter, the FRLM algorithm obtains the robustness to non-Gaussian and Gaussian noises. The performance of the proposed algorithm is validated in simulated acoustic environments. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 3 |
| 2020 | Robust and steerable kronecker product differential beamforming With rectangular microphone arraysabstractDifferential microphone arrays (DMAs), a class of welldesigned small-size arrays combined with differential beamforming, are very useful for processing broadband acoustic, audio, and speech signals in a wide range of applications. However, most efforts in the literature so far have been devoted to linear, circular, and spherical arrays. In this paper, we consider rectangular shapes of planar microphone arrays. Instead of adopting the traditional differential beamforming methods developed in the literature, we present a differential beamforming method based on the so-called Kronecker product. We first decompose the entire rectangular array into two virtual rectangular sub-arrays so that the steering vector of the entire array is the Kronecker product of the steering vectors of the two smaller virtual rectangular sub-arrays. We use the first virtual rectangular array, which is much smaller in size than the entire array but well satisfies the basic requirements for differential beamforming, to design a steerable differential beamformer. For the second virtual rectangular array, we can design either the delay-and-sum (DS) beamformer, which helps to improve the robustness of the global differential beamformer, or an adaptive beamformer, which makes the global differential beamformer adaptive. This method has many interesting properties, particularly the designed beamformer is fully steerable, and its robustness and the array gain can be easily controlled. Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
ICASSP | 2 |
| 2020 | An Improved Solution to the Frequency-Invariant Beamforming with Concentric Circular Microphone ArraysabstractFrequency-invariant beamforming with circular microphone arrays (CMAs) has drawn a significant amount of attention for its steering flexibility and high directivity. However, frequency-invariant beam-forming with CMAs often suffers from the so-called null problem, which is caused by the zeros of the Bessel functions; then, concentric CMAs (CCMAs) are used to deal with this problem. While frequency-invariant beamforming with CCMAs can mitigate the null problem, the beampattern is still suffering from distortion due to s-patial aliasing at high frequencies. In this paper, we find that the spatial aliasing problem is caused by higher-order circular harmonics. To deal with this problem, we take the aliasing harmonics into account and approximate the beampattern with a higher truncation order of the Jacobi-Anger expansion than required. Then, the beam-forming filter is determined by minimizing the errors between the desired directivity pattern and the approximated one. Simulation results show that the developed method can mitigate the distortion of the beampattern caused by spatial aliasing. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2020 | An efficient Kalman filter for the identification of low-rank systems
Laura-Maria Dogariu, Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
Signal Process. | 3 |
| 2020 | A class of multichannel sparse linear prediction algorithms for time delay estimation of speech sources
Hongsen He, Jingdong Chen, Jacob Benesty, Wenxing Zhang, Tao Yang 0039 |
Signal Process. | 3 |
| 2020 | Harmonic beamformers for speech enhancement and dereverberation in the time domain
Jesper Rindom Jensen, Sam Karimian-Azari, Mads Græsbøll Christensen, Jacob Benesty |
Speech Commun. | 4 |
| 2020 | Adaptive and hybrid Kronecker product beamforming for far-field speech signals
Rajib Sharma, Israel Cohen, Jacob Benesty |
Speech Commun. | 3 |
| 2020 | Beamforming With Small-Spacing Microphone Arrays Using Constrained/Generalized LASSOabstractIn this letter, we develop an approach to the design of beamformers with small-spacing uniform linear microphone arrays by incorporating sparseness constraints for attenuating scattered interference incident from some pre-specified ranges of directions of arrival. The design process is formulated as a constrained LASSO problem. By adjusting the value of a tuning parameter, the proposed method can make compromises among three important yet conflicting (especially at low frequencies) performance measures of small-spacing microphone arrays, i.e., the directivity factor (DF), which quantifies the array spatial gain, the white noise gain (WNG), which evaluates the robustness of the beamformer, and the signal-to-interference-ratio (SIR) gain with respect to scattered interference. Simulation results illustrate the properties of the developed approach. Xianghui Wang, Jacob Benesty, Jingdong Chen, Israel Cohen |
IEEE Signal Process. Lett. | 2 |
| 2020 | Joint Sparse Concentric Array Design for Frequency and Rotationally Invariant BeampatternabstractFrequency-invariant concentric arrays are fundamental components in some real-world applications, like teleconferencing, voice service devices, underwater acoustics, and others, where the azimuthal arrival direction of the desired signal is varying. The fact that the demand for limited hardware and computational resources in such applications is essential, motivates the use of a sparse design which can optimize both the number of the required sensors and the complex weights of the beamformer. Herein, we propose a new greedy based joint-sparse design of frequency and rotationally invariant concentric arrays which preserves the properties of the designed directivity pattern for different azimuthal directions of steering. Simulation results show that the greedy sparse design, compared to uniform and random designs, gives superior performance in terms of array gain, and frequency and rotationally invariant beampattern, with a reasonable computational and hardware resources. Yaakov Buchris, Israel Cohen, Jacob Benesty, Alon Amar |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Differential Beamforming on GraphsabstractWe study differential beamforming from a graph perspective. The microphone array used for differential beamforming is viewed as a graph, where its sensors correspond to the nodes, the number of microphones corresponds to the order of the graph, and linear spatial difference equations among microphones are related to graph edges. Specifically, for the first-order differential beamforming with an array of M microphones, each pair of adjacent microphones are directly connected, resulting in M - 1 spatial difference equations. On a graph, each of these equations corresponds to a 2-clique. For the second-order differential beamforming, each three adjacent microphones are directly connected, resulting in M - 2 second-order spatial difference equations, and each of these equations corresponds to a 3-clique. In an analogous manner, the differential microphone array for any order-of-differential beamforming can be viewed as a graph. From this perspective, we then derive a class of differential beamformers, including the maximum white noise gain beamformer, the maximum directivity factor one, and optimal compromising beamformers. Simulations are presented to demonstrate the performance of the derived differential beamformers. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | A Simple Theory and New Method of Differential Beamforming With Uniform Linear Microphone ArraysabstractThis article presents a theoretical study of differential beamforming with uniform linear arrays. By defining a forward spatial difference operator, any order of the spatial difference of the observed signals can be represented as a product of a difference operator matrix and the microphone array observations. Consequently, differential beamforming is implemented in two stages, where the first one obtains spatial difference of the observations and the second stage optimizes the beamformer. The major contributions of this article are as follows. First, we propose a new theory of differential beamforming with uniform linear arrays, which shows clearly the connection between the conventional differential beamforming and the null-constrained differential beamforming methods. This provides some new insight into the design of differential beamformers. Second, we deduce some new differential beamformers, where conventional beamforming may be seen as a particular case. Specifically, we derive the maximum white noise gain (MWNG), maximum directivity factor (MDF), parameterized MDF, and parameterized maximum front-to-back ratio differential beamformers. Third, we further extend the idea of how to design optimal differential beamformers by combining both the observed signals and their spatial differences. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Design of Planar Differential Microphone Arrays With Fractional OrdersabstractDifferential microphone arrays (DMAs) often encounter white noise amplification, especially at low frequencies. If the array geometry and the number of microphones are fixed, one can improve the white noise amplification problem by reducing the DMA order. With the existing differential beamforming methods, the DMA order can only be a positive integer number. Consequently, with a specified beampattern (or a kind of beampattern), reducing this order may easily lead to over compensation of the white noise gain (WNG) and too much reduction of the directivity factor (DF), which is not optimal. To deal with this problem, we present in this article a general approach to the design of DMAs with fractional orders. The major contributions of this article include but are not limited to: 1) we first define a directivity pattern that can achieve a continuous compromise between the pattern corresponding to the maximum DMA order and the omnidirectional pattern; 2) by approximating the beamformer's beampattern with the Jacobi-Anger expansion, we present a method to find the proper differential beamforming filter so that its beampattern matches closely the target directivity pattern of fractional orders; and 3) we show how to determine analytically the proper fractional order of the DMA with a given target beampattern when either the value of the DF or WNG is specified, which is useful in practice to achieve the desired beampattern and spatial gain while maintaining the robustness of the DMA system. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | A Recursive Least-squares Algorithm Based on the Nearest Kronecker Product DecompositionabstractThe recursive least-squares (RLS) adaptive filter is an appealing choice in system identification problems, mainly due to its fast convergence rate. However, this algorithm is computationally very complex, which may make it useless for the identification of high length impulse responses, like in echo cancellation. In this paper, we focus on a new approach to improve the efficiency of the RLS algorithm. The basic idea is to exploit the impulse response decomposition based on the nearest Kronecker product and low-rank approximation. Thus, a high-dimension system identification problem is reformulated in terms of low-dimension problems, which are tensorized together. Simulations performed in the context of echo cancellation indicate the good performance of the RLS algorithm based on this approach. Camelia Elisei-Iliescu, Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
ICASSP | 3 |
| 2019 | Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone ArraysabstractSmall aperture circular microphone arrays (CMAs) have been widely used in many applications such as teleconferencing, smartspeakers, and robotics. A critical component of such arrays is the differential beamformer, which can achieve relatively high spatial gains with the same beampatterns at most frequencies. Among different differential beamforming approaches that were developed in the literature, the minimum-norm one has attracted much interest as it can deal better with sensors' self noise, sensor mismatch, and beamformer's irregularity at some frequencies due to the zeros of the Bessel functions. In our previous study, we have investigated the performance of the minimum-norm differential beamformer with uniform CMAs (UCMAs) in the 2-dimensional (2D) space where the sound sources and the sensors are assumed to be in the same plane. But in practice, this assumption is generally not true. So, in this paper, we investigate the properties and limitations of the minimum-norm differential beamformer in the 3-dimensional(3D) space. Through theoretical study as well as simulations, we show that the minimum-norm differential beamformer is effective in dealing with the problem of white noise amplification and irregularity of the beampatterns and the directivity factor (DF) if the steering angles are within or near the sensor plane, but it becomes less and less effective as the beamformer is steered away from this plane. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2019 | Design of Optimal Linear Differential Microphone Arrays Based Array Geometry OptimizationabstractThis paper presents a method to design optimal linear differential microphone arrays (DMAs) by optimizing the array geometry. By constraining the DMA beamformer to achieve a given target value of the directivity factor (DF) with a specified target frequency-invariant beampattern while achieving also the highest possible white noise gain (WNG), an optimization algorithm is developed, which consists of the following two steps. 1) The full frequency band of interest is divided into a few subbands. At every subband, the entire linear array is divided into subarrays and the number of subarrays depends on the total number of the sensors and the order of the DMA. A cost function is then defined, which is minimized to determine what subarray produces the optimal performance. 2) The subband optimal subarrays are then combined across the entire frequency band to form a fullband cost function, from which the geometry of the entire array is optimized. These two steps are repeated with the particle swarm optimization (PSO) algorithm until the desired array performance is reached. Simulation results demonstrate that the proposed method can obtain the target DF with a frequency-invariant beampattern over a wide band of frequencies while maintaining a reasonable level of WNG. Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2019 | On the Design of Flexible Kronecker Product Beamformers with Linear Microphone ArraysabstractThis paper proposes a method for the design of flexible Kronecker product beamformers based on the decomposition of the steering vector of a physical array as a Kronecker product of steering vectors of two smaller virtual arrays. With this decomposition, the global beamforming filter is designed by optimizing the two sub-beamformers in a cascaded manner, which can offer much flexibility to control the performance of beamforming or control the compromise between different, conflicted performance measures. In comparison with a recently developed method that restricts the number of microphones of the given physical array to a multiplication of two integers, each corresponding to the number of sensors of one virtual array, the approach in this work decomposes the physical array in such a way that the sensors in the two virtual arrays may share positions and the number of microphones of the physical array can be any positive integer. Simulations demonstrate the properties of the proposed approach. Wenxing Yang, Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
ICASSP | 3 |
| 2019 | Nonlinear Kronecker product filtering for multichannel noise reduction
Gal Itzhak, Jacob Benesty, Israel Cohen |
Speech Commun. | 2 |
| 2019 | Incoherent Synthesis of Sparse Arrays for Frequency-Invariant BeamformingabstractFrequency-invariant beamformers are used to prevent signal waveform distortions in real world applications like audio, underwater acoustics, and radar. Most of existing methods assume uniform arrays, and only few consider sparse designs, which may lead to higher performance in terms of robustness and directivity factor. We propose an incoherent approach that first determines for each frequency bin a sparse set of sensors positions. Subsequently, by using tools of dimensionality reduction and clustering, these selections are merged together yielding the optimal sensors on a sparse array layout. We present design examples of sparse linear and planar superdirective array designs. We show that the proposed incoherent sparse design obtains superior performance in terms of white noise gain, directivity factor, and computational load compared to a uniform array design and compared to a coherent sparse approach, where the sensors' locations and the beamformer coefficients are optimized simultaneously for all frequencies. Yaakov Buchris, Alon Amar, Jacob Benesty, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Differential Kronecker Product BeamformingabstractDifferential beamformers have attracted much interest over the past few decades. In this paper, we introduce differential Kronecker product beamformers that exploit the structure of the steering vector to perform beamforming differently from the well-known and studied conventional approach. We consider a class of microphone arrays that enable to decompose the steering vector as a Kronecker product of two steering vectors of smaller virtual arrays. In the proposed approach, instead of directly designing the differential beamformer, we break it down following the decomposition of the steering vector, and show how to derive differential beamformers using the Kronecker product formulation. As demonstrated, the Kronecker product decomposition facilitates further flexibility in the design of differential beamformers and in the tradeoff control between the directivity factor and the white noise gain. Israel Cohen, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Recursive Least-Squares Algorithms for the Identification of Low-Rank SystemsabstractThe recursive least-squares (RLS) adaptive filter is an appealing choice in many system identification problems. The main reason behind its popularity is its fast convergence rate. However, this algorithm is computationally very complex, which may make it useless for the identification of long length impulse responses, like in echo cancellation. Computationally efficient versions of the RLS algorithm, like those based on the dichotomous coordinate descent (DCD) iterations or QR decomposition techniques, reduce the complexity, but still have to face the challenges related to long length adaptive filters (e.g., convergence/tracking capabilities). In this paper, we focus on a different approach to improve the efficiency of the RLS algorithm. The basic idea is to exploit the impulse response decomposition based on the nearest Kronecker product and low-rank approximation. In other words, a high-dimension system identification problem is reformulated in terms of low-dimension problems, which are combined together. This approach was recently addressed in terms of the Wiener filter, showing appealing features for the identification of low-rank systems, like real-world echo paths. In this paper, besides the development of the RLS algorithm based on this approach, we also propose a variable regularized version of this algorithm (using the DCD method to reduce the complexity), with improved robustness to double-talk. Simulations are performed in the context of echo cancellation and the results indicate the good performance of these algorithms. Camelia Elisei-Iliescu, Constantin Paleologu, Jacob Benesty, Cristian Lucian Stanciu, Cristian Anghel, Silviu Ciochina |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | On the Design of Target Beampatterns for Differential Microphone ArraysabstractDifferential microphone arrays (DMAs) have many interesting properties and have been widely used in acoustic, audio, and speech applications. A critical part of a DMA is the differential beamformer, which is generally designed in two important steps: 1) specifying a target beampattern based on what differential sound pressure field the DMA is expected to respond to and 2) designing the differential beamforming filter so that the resulting beampattern matches the target one. Most efforts in the study of DMAs so far have focused on the second step while choosing one of the limited patterns available in the literature as the target beampattern. Since it governs how the array performs, how to design the target beampattern is an important problem, which this paper addresses. The major contributions of this paper consists of the following four aspects. First, a positive superposition theorem is presented, which shows that the linear combination of effective beampatterns with non-negative coefficients is always an effective beampattern. Second, we propose a general approach to the design of target DMA beampatterns based on the positive superposition theorem. Third, an overview of the classical target beampatterns is provided and discussion is made on how to form effective base patterns. Fourth, we show that the smallest first null of a DMA is π(2N) with N being the DMA order, which provides the rule of setting nulls in practice. Finally, with examples, we show that with the use of the alternating-direction-method-of-multipliers algorithm, the proposed approach is able to generate useful DMA target beampatterns. Chao Pan 0001, Jingdong Chen, Jacob Benesty, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | On Robust and High Directive Beamforming With Small-Spacing Microphone Arrays for Scattered SourcesabstractThis paper is devoted to beamforming with small-spacing microphone arrays for processing broadband and scattered acoustic sources. It presents a maximum diffuse noise gain (MDNG) beamformer in this context using the joint diagonalization technique, which is effective in suppressing diffuse and directional noise, but at a price of low white noise gain (WNG). We also introduce a maximum WNG (MWNG) beamformer, which is robust to the array imperfections, but paying a price of sacrificing the diffuse noise gain (DNG). To make a tradeoff between WNG and DNG so that the beamformer, on the one hand, can achieve high directivity and, on the other hand, is robust to implement, we propose a generalized MDNG beamformer, which includes both the MDNG and MWNG beamformers as particular cases. Simulations are conducted to illustrate the properties and advantages of the proposed beamformers. Xianghui Wang, Israel Cohen, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Identification of Bilinear Forms with the Kalman FilterabstractIn this paper, we develop the Kalman filter for the identification of bilinear forms. In this framework, the bilinear term is defined with respect to the impulse responses of a spatiotemporal model, which resembles a multiple-input/single-output system. Recently, the identification of such bilinear forms was addressed in terms of the Wiener filter and conventional adaptive algorithms, i.e., least-mean-square and recursive least-squares. In this work, apart from the derivation of the Kalman filter tailored for the identification of bilinear forms, a simplified (i.e., low complexity) version of the algorithm is also presented. Simulation results support the theoretical findings and indicate the good performance of the proposed solutions. Laura-Maria Dogariu, Constantin Paleologu, Silviu Ciochina, Jacob Benesty, Pablo Piantanida |
ICASSP | 4 |
| 2018 | On the Design of Robust Steerable Frequency-Invariant Beampatterns with Concentric Circular Microphone ArraysabstractThis paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs). We develop a beamforming algorithm based on an optimal approximation of the beamformer's beampattern with the Jacobi-Anger expansion. In comparison with the existing frequency-invariant beamformers with either circular microphone arrays (CMAs) or CCMAs, the developed algorithm offers the following advantages: 1) it can mitigate the deep-null problem encountered in CMAs and therefore has a consistent directivity factor over the frequency range of speech signals; 2) it is more flexible in terms of steering flexibility and the resulting beampattern can be steered to any direction; and 3) it does not require the microphones in different rings of the CCMA to be aligned, which is very useful in practice, particularly when microphone arrays with small and compact apertures have to be used. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2018 | On Speech Enhancement Using Microphone Arrays in the Presence of Co-Directional InterferenceabstractBeamforming using microphone arrays has been widely used for enhancing speech signals of interest and suppressing noise and interference in a wide range of applications. In order to make it work, beamforming generally assumes that the speech source of interest and the interference source are incident to the array from different directions. In this paper, we study the case where both the speech and interference sources come from the same direction. A linearly constrained minimum variance (LCMV) beamformer is derived in this scenario based on the so-called widely linear (WL) estimation framework in the frequency domain. We analyze this beamformer and show how its performance depends on the second-order non-circularity of the desired speech and interference sources. Xin Leng, Jingdong Chen, Jacob Benesty, Israel Cohen |
ICASSP | 3 |
| 2018 | A Single-Channel Noise Reduction Filtering/Smoothing Technique in the Time DomainabstractIn this paper, we present a single-channel smoothing-and-filtering technique for noise reduction in the time domain. Unlike traditional noise reduction methods, which directly apply a noise reduction filter to the noisy signal, the developed technique achieves noise reduction in two steps. It first applies a time smoothing window to the noisy signal, which, on the one hand, can help reduce high frequency noise and, on the other hand, can help leverage the correlation between successive signal samples. A noise reduction filter is then applied to the smoothed noisy signal to estimate the speech signal of interest. Three optimal and suboptimal noise reduction filters are derived, including the Wiener, maximum signal-to-noise-ratio (SNR), and tradeoff filters. Simulation results reveal that the developed method can produce better noise reduction performance, i.e., higher gains in the perceptual-evaluation-of-speech-quality (PESQ) score, than the traditional methods without smoothing. Ningning Pan, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2018 | Frequency-Domain Design of Asymmetric Circular Differential Microphone ArraysabstractCircular differential microphone arrays (CDMAs) facilitate compact superdirective beamformers whose beampatterns are nearly frequency invariant. In contrast to linear differential microphone arrays where the optimal steering direction is at the endfire, CDMAs provide perfect steering for all azimuthal directions. Herein, we extend the traditional symmetric model of DMAs and establish an analytical asymmetric model for Nth-order CDMAs. This model exploits the circular geometry to eliminate the inherent limitation of symmetric beampatterns associated with a linear geometry and allows also asymmetric beampatterns. This new model is then used to develop asymmetric versions of two optimal commonly used beampatterns namely the hypercardioid and the supercardioid. Experimental results demonstrate the advantages of the asymmetric model compared to the traditional symmetric one, when additional directional constraints are imposed. The proposed model yields superior performance in terms of white noise gain, directivity factor, and front-to-back ratio, as well as more flexible design of nulls for the interfering signals. Yaakov Buchris, Israel Cohen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Noise Robust Frequency-Domain Adaptive Blind Multichannel Identification With ℓp-Norm ConstraintabstractBlind multichannel identification is a challenging problem in many domains. The normalized multichannel frequency-domain least-mean-square (NMCFLMS) algorithm was developed to blindly identify a single-input multiple-output acoustic system, which can yield good performance in noise-free environments. However, the robustness of this algorithm to noise has been shown to be problematic. One way to improve the robustness is by applying a constraint on the spectral flatness of the channel impulse responses, which led to the development of the so-called robust normalized multichannel frequency-domain least-mean-square (RNMCFLMS) algorithm. This spectral flatness constraint, however, may not be always proper or reasonable in realistic acoustic environments. In this paper, we develop an ℓp-norm constraint based robust normalized multichannel frequency-domain least-mean-square (ℓp-RNMCFLMS) algorithm. The ℓp-norm constraint is introduced into the NMCFLMS algorithm to control the effect of different ℓp-norm penalties on the adaptive filter for the impulse responses with different degrees of sparseness. Numerical and realistic experiments justify the effectiveness of the proposed ℓp-RNMCFLMS algorithm. Hongsen He, Jingdong Chen, Jacob Benesty, Tao Yang 0039 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Insights Into Frequency-Invariant Beamforming With Concentric Circular Microphone ArraysabstractThis paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs) and presents an approach to the design of frequency-invariant and symmetric beampatterns. We first apply the Jacobi-Anger expansion to each ring of the CCMA to approximate the beampattern. The beamformer is then designed by using all the expansions from different rings. In comparison with the existing work in the literature where a Jacobi-Anger expansion of the same order is applied to different rings, here in this contribution the order of the Jacobi-Anger expansion at a ring is related to its number of sensors and, as a result, the expansion order at different rings may be different. The developed approach is rather general. It is not only able to mitigate the deep nulls problem in the directivity factor and the white noise gain, that is common to circular microphone arrays (CMAs), and improve the steering flexibility, but is also flexible to use in practice where a smaller ring can have less microphones than a larger one. We discuss the conditions for the design ofNth-order symmetric beampatterns and examples of frequency-invariant beampatterns with commonly used array geometries such as CMAs, CMAs with a sensor at the center, and CCMAs. We show the advantage of adding one microphone at the center of either a CMA or a CCMA, i.e., circumventing the deep nulls problem caused by the 0th-order Bessel function. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Linear System Identification Based on a Kronecker Product DecompositionabstractLinear system identification is a key problem in many important applications, among which echo cancelation is a very challenging one. Due to the long length impulse responses (i.e., echo paths) to be identified, there is always room (and needs) to improve the performance of the echo cancelers, especially in terms of complexity, convergence rate, robustness, and accuracy. In this paper, we propose a new way to address the system identification problem (from the echo cancelation perspective), by exploiting an optimal approximation of the impulse response based on the nearest Kronecker product decomposition. Also, we make a first step toward this direction, by developing an iterative Wiener filter based on this approach. As compared to the conventional Wiener filter, the proposed solution is much more attractive since its gain is twofold. First, the matrices to be inverted (or, preferably, linear systems to be solved) are smaller as compared to the conventional approach. Second, as a consequence, the iterative Wiener filter leads to a good estimate of the impulse response, even when a small amount of data is available for the estimation of the statistics. Simulation results support the theoretical findings and indicate the good results of the proposed approach, for the identification of different network and acoustic impulse responses. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Robust multichannel TDOA estimation for speaker localization using the impulsive characteristics of speech spectrumabstractTime delay estimation (TDE) plays an important role in localizing and tracking radiating acoustic sources. Although many efforts have been devoted to this problem in the literature, the robustness of TDE with respect to noise and reverberation remains a great challenge for practical systems. In this paper, we investigate the TDE problem in acoustic single-input/multiple-output (SIMO) systems in reverberant and noisy environments. We first define a Cauchy estimator in the frequency domain, which is robust in dealing with speech as the SIMO system's excitation. This robust estimator is then used to construct a cost function, from which a robust multichannel frequency-domain adaptive filter is deduced. This adaptive algorithm is subsequently employed to blindly identify the acoustic impulse responses between the source and the microphones. Finally, the time difference of arrival is determined from the identified channel responses. Hongsen He, Jingdong Chen, Jacob Benesty, Yingyue Zhou, Tao Yang 0039 |
ICASSP | 3 |
| 2017 | Study of the frequency-domain multichannel noise reduction problem with the householder transformationabstractThis paper presents an approach to the multichannel noise reduction problem. It first transforms the multichannel noisy speech signals into the frequency domain. A Householder transformation is then constructed, which converts the multichannel coefficients in each frequency bin into two components: one dominated by speech and the other dominated by noise. A Wiener filter is subsequently formed to achieve an estimate of the noise in the speech dominated component from the noise dominated component. The enhanced speech is then obtained by subtracting the noise estimate from the speech dominated component. This approach consists of two critical steps: construction of the Householder transformation and formation of the noise reduction Wiener filter. If the source incidence angle is known a priori, the Householder transformation can be directly constructed using the steering vector and the optimal estimate of the signal of interest can then be obtained by applying the Wiener filter. If the source incidence angle is not known a priori, the Householder transformation can be constructed from a hypothesized incidence angle. Then, the optimal signal estimate is obtained by searching the maximum of the variance of the enhanced signal with the Wiener filter in the interested range of the incidence angle. Gongping Huang, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2017 | Distributed max-SINR speech enhancement with ad hoc microphone arraysabstractIn recent years, signal processing with ad hoc microphone arrays has attracted a lot of attention. Speech enhancement in noisy, interfered, and reverberant environments is one of the problems targeted by ad hoc microphone arrays. Most of the proposed solutions require knowledge of fingerprints, such as acoustic transfer functions, which may not be known as accurately as required in practical situations. In this paper, a distributed signal subspace filtering method is proposed which is not restricted to a special graph topology. Here, the maximum signal to interference-plus-noise ratio (max-SINR) criterion is used with the primal-dual method of multipliers for distributed filtering. The paper investigates the convergence of the algorithm in both synchronous and asynchronous schemes, and also discusses some practical pros and cons. The applicability of the proposed method is demonstrated by means of simulation results. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Richard Heusdens, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 4 |
| 2017 | A minimum variance partially distortionless response filter for single-channel noise reductionabstractThis paper deals with the problem of single-channel noise reduction. Thanks to the eigenvalue decomposition, we arrange the eigenvalues of the speech correlation matrix in such a way that all the spectral mode signal-to-noise ratios (SNRs) of the noisy speech are ordered in a descending manner. By maintaining no speech distortion in the spectral modes with high input SNRs while allowing some degree of speech distortion in the modes with low input SNRs, we develop a minimum variance partially distortionless response (MVPDR) filter. We first formulate the problem and derive this filter within the general filtering framework. Then, the MVPDR filter is applied to the single-channel noise reduction problem in both the time and time-frequency domains. In comparison with the minimum variance distortionless response (MVDR) filter based on the subspace decomposition, the developed MVPDR filter can provide much more freedom for controlling the compromise between noise reduction and speech distortion to achieve higher speech quality. Simulations are conducted and preliminary results justify the advantages of the deduced MVPDR filter. Xianghui Wang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2017 | On the Identification of Bilinear Forms With the Wiener FilterabstractIn this letter, the identification problem of bilinear forms with the Wiener filter is addressed. The contribution is twofold. First, a different approach is introduced, by defining the bilinear term with respect to the impulse responses of a spatiotemporal model, in the context of multiple-input/single-output systems. Second, two versions of the Wiener filter (namely direct and iterative) are developed in this context. Moreover, the advantage of the iterative Wiener filter is outlined as compared to the direct solution. The results of the simulations, which are performed from a system identification perspective, support the theoretical findings. Jacob Benesty, Constantin Paleologu, Silviu Ciochina |
IEEE Signal Process. Lett. | 1 |
| 2017 | On the Design of Frequency-Invariant Beampatterns With Uniform Circular Microphone ArraysabstractThis paper deals with two critical issues about uniform circular arrays (UCAs): frequency-invariant response and steering flexibility. It focuses on some optimal design of frequency-invariant beampatterns in any desired direction along the sensor plane. The major contributions are as follows. 1) We explain how to include the steering information in the desired directivity pattern. 2) We show that the optimal approximation of the beamformer's beampattern with a UCA from a least-squares error perspective is the Jacobi-Anger expansion. 3) We develop an approach to the design of any desired symmetric directivity pattern, where the deduced beampattern is almost frequency invariant and its main beam can be pointed to any wanted direction in the sensor plane. 4) With the proposed approach, we derive an explicit form of the white noise gain (WNG) and the directivity factor (DF), and explain clearly the white noise amplification problem at low frequencies and the DF degradation at high frequencies. The analysis also indicates that increasing the number of microphones can always improve the WNG. We show that the proposed method is a generalization of circular differential microphone arrays. The relationship between the proposed method and the so-called circular harmonics beamformers is also discussed. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | On time delay estimation based on multichannel spatiotemporal sparse linear predictionabstractNoise and reverberation can significantly affect the performance of time delay estimation (TDE) in room acoustic environments. The multichannel cross-correlation coefficient (MCCC) algorithm, which extends the traditional cross-correlation method from two to multiple channels, can exploit the spatial information among multiple microphones to improve the robustness of TDE with respect to environmental noise; but this algorithm is not robust to reverberation. The multichannel spatiotemporal prediction (MCSTP) algorithm uses both the spatial and temporal information provided by the array. This algorithm improves significantly the robustness of TDE with respect to reverberation; however, it is found sensitive to noise. In this paper, we develop a multichannel spatiotemporal sparse prediction (MCSTSP) algorithm for TDE. This algorithm obtains a good compromise between robustness of TDE to noise and that to reverberation through making a tradeoff between pre-whitening and non-prewhitening. This is achieved via adjusting a regularization parameter, which is solved by an augmented Lagrangian alternating direction method of multipliers (ADMM). The property of this developed algorithm is justified with numerical experiments in both noisy and reverberant environments. Hongsen He, Jingdong Chen, Jacob Benesty, Tao Yang 0039 |
ICASSP | 3 |
| 2016 | Variable span filters for speech enhancementabstractIn this work, we consider enhancement of multichannel speech recordings. Linear filtering and subspace approaches have been considered previously for solving the problem. The current linear filtering methods, although many variants exist, have limited control of noise reduction and speech distortion. Subspace approaches, on the other hand, can potentially yield better control by filtering in the eigen-domain, but traditionally these approaches have not been optimized explicitly for traditional noise reduction and signal distortion measures. Herein, we combine these approaches by deriving optimal filters using a joint diagonalization as a basis. This gives excellent control over the performance, as we can optimize for noise reduction or signal distortion performance. Results from real data experiments show that the proposed variable span filters can achieve better performance than existing filters. In terms of output SNR, the gain was more than 8 dB, and more than 0.1 in mean opinion score in the conducted experiments. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 2 |
| 2016 | Subspace superdirective beamformers based on joint diagonalizationabstractAlthough they have been intensively studied and used in many applications due to their high directivity factor (DF), superdirective beamformers are sensitive to sensor noise and mismatch between sensors. This paper studies the problem of superdirective beamforming combined with the joint diagonalization method. We develop a subspace superdirective beamforming approach, which can achieve a good compromise between a high DF and white noise amplification. Simulations are performed to justify our theoretical analysis and demonstrate the good properties of this subspace superdirective beamforming approach. Changlei Li, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 2 |
| 2016 | A partitioned approach to signal separation with microphone ad hoc arraysabstractIn this paper, a blind algorithm is proposed for speech enhancement in multi-speaker scenarios, in which interference rejection is the main objective. Here, the ad hoc array is broken into microphone duples which are used to partition the array into local sub-arrays. The core algorithm takes advantage of differences in signal structure in each duple. A geometric mean filter is then used to merge the output signals obtained with different duples, and to form a global broadband maximum signal-to-interference ratio (SIR) enhancement apparatus. The resulting filter outputs are enhanced acoustic signals in terms of SIR, as shown with experiments. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen |
ICASSP | 3 |
| 2016 | A single-channel noise cancelation filter in the short-time-fourier-transform domainabstractThis paper develops a single-channel noise cancelation filter in the short-time Fourier transform (STFT) domain by combining the subspace method and the optimal filtering technique via joint diagonalization of the desired clean speech and noise signal correlation matrices. This filter is shown to be flexible in controlling the compromise between the output signal-to-noise ratio (oSNR) and the amount of speech distortion. Simulations are performed to justify the property of this filter. Xianghui Wang, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2016 | An optimized NLMS algorithm for system identification
Silviu Ciochina, Constantin Paleologu, Jacob Benesty |
Signal Process. | 3 |
| 2016 | Single-channel noise reduction via semi-orthogonal transformations and reduced-rank filtering
Jacob Benesty, Jingdong Chen |
Speech Commun. | 2 |
| 2016 | Superdirective Beamforming Based on the Krylov MatrixabstractSuperdirective beamforming has attracted a significant amount of research interest in speech and audio applications, since it can maximize the directivity factor (DF) given an array geometry and, therefore, is efficient in dealing with signal acquisition in diffuse-like noise environments. However, this beamformer is very sensitive to sensor self-noise and mismatch among sensors, which considerably restricts its use in practical systems. This paper develops an approach to superdirective beamforming based on the Krylov matrix. We show that the columns of a proposed Krylov matrix, which span a chosen dimension of the whole space, are interesting beamformers; consequently, all different linear combinations of those columns lead to beamformers that have good properties. In particular, we develop the Krylov maximum white noise gain and Krylov maximum DF beamformers, which are obtained by maximizing the WNG and the DF, respectively. By properly choosing the dimension of the Krylov subspace, the developed beamformers that can make a compromise between reasonable values of the DF and white noise amplification. We also extend the basic idea to the design of the Krylov maximum front-to-back ratio, parametric superdirective, and parametric supercardioid beamformers. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Noise Reduction with Optimal Variable Span Linear FiltersabstractIn this paper, the problem of noise reduction is addressed as a linear filtering problem in a novel way by using concepts from subspace-based enhancement methods, resulting in variable span linear filters. This is done by forming the filter coefficients as linear combinations of a number of eigenvectors stemming from a joint diagonalization of the covariance matrices of the signal of interest and the noise. The resulting filters are flexible in that it is possible to trade off distortion of the desired signal for improved noise reduction. This tradeoff is controlled by the number of eigenvectors included in forming the filter. Using these concepts, a number of different filter designs are considered, like minimum distortion, Wiener, maximum SNR, and tradeoff filters. Interestingly, all these can be expressed as special cases of variable span filters. We also derive expressions for the speech distortion and noise reduction of the various filter designs. Moreover, we consider an alternative approach, wherein the filter is designed for extracting an estimate of the noise signal, which can then be extracted from the observed signals, which is referred to as the indirect approach. Simulations demonstrate the advantages and properties of the variable span filter designs, and their potential performance gain compared to widely used speech enhancement methods. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Design of Directivity Patterns with a Unique Null of Maximum MultiplicityabstractDifferential beamforming is one of the most popular beamforming approaches, which has the great potential to form frequency-invariant directivity patterns. In this paper, we study the design of beampatterns with multiple nulls in the same direction, which is clearly different from the design of beampatterns with distinct nulls. Our contributions are as follows. First, we show how to constrain multiple nulls to the same direction and design the desired beampattern with both the traditional and robust approaches. Second, we derive an explicit form of the white noise gain (WNG) of the traditional approach as a function of the frequency, interelement spacing, and null direction, which shows that the cardioid is the optimal beampattern as far as the WNG is concerned. Third, we prove that the WNG improvement of the robust approach rarely depends on the null direction at low frequencies. Finally, considering the fact that the robust differential beamforming approach may produce a frequency-dependent beampattern while improving the WNG, we develop a weighted-norm approach that can make a good compromise between the robustness of differential beamforming with respect to white noise and the frequency-invariant beampattern. The performance of the developed approach is verified by simulations. Chao Pan 0001, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Reduced-Order Robust Superdirective Beamforming With Uniform Linear Microphone ArraysabstractSensor arrays for audio and speech signal acquisition are generally required to have frequency-invariant beampatterns to avoid adding spectral distortion to the broadband signals of interest. One way to obtain frequency-invariant beampatterns is via superdirective beamforming. However, traditional superdirective beamformers may cause significant white noise amplification (particularly at low frequencies), making them sensitive to uncorrelated white noise. To circumvent the problem of white noise amplification, a method was developed to find the superdirective beamforming filter with a constraint on the white noise gain (WNG), leading to the so-called WNG-constrained superdirective beamformer. But this method damages the frequency invariance of the beampattern. In this paper, we develop a flatness-constrained robust superdirective beamformer. We divide the overall beamformer into two subbeamformers, which are convolved together: one subbeamformer forms a lower order superdirective beampattern while the other attempts to improve the WNG. We show that this robust approach can improve the WNG while limiting the frequency dependency of the beampattern at the same time. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A Framework for Speech Enhancement With Ad Hoc Microphone ArraysabstractSpeech enhancement is vital for improved listening practices. Ad hoc microphone arrays are promising assets for this purpose. Most well-established enhancement techniques with conventional arrays can be adapted into ad hoc scenarios. Despite recent efforts to introduce various ad hoc speech enhancement apparatus, a common framework for integration of conventional methods into this new scheme is still missing. This paper establishes such an abstraction based on inter and intra subarray speech coherencies. Along with measures for signal quality at the input of subarrays, a measure of coherency is proposed both for subarray selection in local enhancement approaches, and also for selecting a proper global reference when more than one subarray are used. Proposed methods within this framework are evaluated with regard to quantitative and qualitative measures, including array gains, the speech distortion ratio, the PESQ measure, and the STOI intelligibility measure. Major findings in this work are the observed changes in the superiority of different methods for certain conditions. When perceptual quality or intelligibility of the speech are the ultimate goals, there are turning points where the MVDR and the LCMV are superior to Wiener-based methods. Also, for certain scenarios, local approaches may be preferred to global ones. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Investigation of a parametric gain approach to single-channel speech enhancementabstractThis paper investigates a parametric gain approach to single-channel noise reduction in the frequency domain. In comparison with the traditional parametric Wiener gain, the major novelty of this presented approach is that the parametric gain is formulated to estimate the noise by using the mean-squared error (MSE) between the noise and the noise estimate. The enhanced signal is then obtained by subtracting the noise estimate from the noisy observation signal. We show that this new method is more practical to implement and can produce better noise reduction performance as compared to the traditional parametric Wiener filtering techniques if the order of the parametric gain is not equal to 1. If the order is 1, the parametric gain is similar to the traditional Wiener gain. Simulation results are presented to illustrate the properties of this new approach. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2015 | Pseudo-coherence-based MVDR beamformer for speech enhancement with ad hoc microphone arraysabstractSpeech enhancement with distributed arrays has been met with various methods. On the one hand, data independent methods require information about the position of sensors, so they are not suitable for dynamic geometries. On the other hand, Wiener-based methods cannot assure a distortionless output. This paper proposes minimum variance distortionless response filtering based on multichannel pseudo-coherence for speech enhancement with ad hoc microphone arrays. This method requires neither position information nor control of the trade-off used in the distortion weighted methods. Furthermore, certain performance criteria are derived in terms of the pseudo-coherence vector, and the method is compared with the multichannel Wiener filter. Evaluation shows the suitability of the proposed method in terms of noise reduction with minimum distortion in ad hoc scenarios. Vincent Mohammad Tavakoli, Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
ICASSP | 4 |
| 2015 | Optimal single-channel noise reduction filtering matrices from the pearson correlation coefficient perspectiveabstractThis paper studies the problem of single-channel noise reduction in the time domain, where an estimate of a vector of the desired clean speech is achieved by filtering a frame of the noisy signal with a rectangular filtering matrix. The core issue with this problem formulation is then the estimation of the optimal filtering matrix. The squared Pearson correlation coefficient (SPCC) is used. We show that different optimal filtering matrices can be derived by maximizing or minimizing the SPCCs between different signals. For example, maximizing the SPCC between the enhanced signal and the filtered speech gives the reduced-rankWiener and minimum distortion (MD) filtering matrices while minimizing the SPCC gives the minimum noise (MN) and another reduced-rank Wiener filtering matrices. Simulation results are presented to illustrate the properties of these filtering matrices. Jiaolong Yu, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 2 |
| 2015 | Optimal design of directivity patterns for endfire linear microphone arraysabstractDirectivity pattern or beampattern is an important performance measure in all fixed beamformers. Given a microphone array, how to design the beamforming filter so that the resulting directivity pattern is close to the desired one is a critical issue. In this paper, we study the design of such patterns for endfire uniform linear microphone arrays. By considering the frequency-independent Chebyshev pattern as the desired one, we derive an optimal beamforming filter based on the minimization of the mean-squared error (MSE) under the distortionless constraint. It is shown that the proposed beamformer design can generate beampatterns that are very close to the desired ones and, the larger is the number of microphones, the better is the designed beampattern. Liheng Zhao, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2015 | Combined Beamformers for Robust Broadband Regularized Superdirective BeamformingabstractSuperdirective fixed beamformers are known to attain high directivity factors, but are extremely sensitive to uncorrelated noise and slight errors in the array elements, which are modeled by the beamformer white noise gain measure. The delay-and-sum beamformer, on the other hand, manages to maximize the white noise gain, but suffers from a very low directivity factor. In this paper, we discuss the design of a broadband beamformer which controls both the directivity factor and the white noise gain. We combine a regularized version of the superdirective beamformer together with the delay-and-sum beamformer to create a robust regularized superdirective beamformer. We derive analytic closed-form expressions of the beamformer gain responses, and extend them to derive a beamformer with full control of the desired white noise gain or the directivity factor. The proposed approach offers a simple and robust broadband beamformer with controllable characteristics, shown here through persuasive simulation results. Reuven Berkun, Israel Cohen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Joint Spatio-Temporal Filtering Methods for DOA and Fundamental Frequency EstimationabstractIn this paper, spatio-temporal filtering methods are proposed for estimating the direction-of-arrival (DOA) and fundamental frequency of periodic signals, like those produced by the speech production system and many musical instruments using microphone arrays. This topic has quite recently received some attention in the community and is quite promising for several applications. The proposed methods are based on optimal, adaptive filters that leave the desired signal, having a certain DOA and fundamental frequency, undistorted and suppress everything else. The filtering methods simultaneously operate in space and time, whereby it is possible resolve cases that are otherwise problematic for pitch estimators or DOA estimators based on beamforming. Several special cases and improvements are considered, including a method for estimating the covariance matrix based on the recently proposed iterative adaptive approach (IAA). Experiments demonstrate the improved performance of the proposed methods under adverse conditions compared to the state of the art using both synthetic signals and real signals, as well as illustrate the properties of the methods and the filters. Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty, Søren Holdt Jensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Theoretical Analysis of Differential Microphone Array Beamforming and an Improved SolutionabstractDifferential microphone arrays (DMAs), which are responsive to the differential sound pressure field, have attracted much attention due to their properties of frequency-invariant beampatterns, small apertures, and potential of maximum directivity. Traditionally, DMAs are designed and implemented in a multistage (cascade) way, where a proper time delay is used in each stage to form a beampattern of interest. Recently, it was reported that DMAs can be designed by solving a linear system of equations formed from the information about the nulls of the desired beampattern. This paper deals with the problem of beamforming with linear DMAs. Its major contributions are as follows. 1) By using the spatial${\cal Z}$transform, we present some theoretical analysis of both the traditional cascade and new null-constrained DMA beamforming. It is shown that the cascade and null-constrained DMAs of the same order with the same number of sensors are theoretically identical. 2) We develop a two-stage approach to the study of the robust DMA beamformer, which is based on the principle of maximizing the white noise gain (WNG). The first-stage of this approach is in the structure of the traditional non-robust DMA while the second-stage filter is optimized for improving the WNG. 3) Using the two-stage approach, we show that the robust DMA beamformer may introduce extra nulls in the beampattern at high frequencies; particularly, it introduces$M - N - 1$extra nulls if the interelement spacing is equal to half of the wavelength, where$M$and$N$are the number of sensors and the DMA order, respectively. 4) We develop a method that can solve the extra-null problem while maximizing the WNG in robust DMA beamforming, i.e., a robust solution with a frequency-invariant beampattern. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Noise reduction in the time domain using joint diagonalizationabstractA new filter design based on joint diagonalization of the clean speech and noise covariance matrices is proposed. First, an estimate of the noise is found by filtering the observed signal. The filter for this is generated by a weighted sum of the eigenvectors from the joint diagonalization. Second, an estimate of the desired signal is found by subtraction of the noise estimate from the observed signal. The filter can be designed to obtain a desired trade-off between noise reduction and signal distortion, depending on the number of eigenvectors included in the filter design. This is explored through simulations using a speech signal corrupted by car noise, and the results confirm that the output signal-to-noise ratio and speech distortion index both increase when more eigenvectors are included in the filter design. Sidsel Marie Nørholm, Jacob Benesty, Jesper Rindom Jensen, Mads Græsbøll Christensen |
ICASSP | 2 |
| 2014 | A Kalman filter with individual control factors for echo cancellationabstractIn echo cancellation, the main goal is to recover the near-end signal from the error signal of the adaptive filter, which identifies the echo path. In this context, the Kalman filter represents a very appealing choice, since its basic criterion follows the minimization of the system misalignment (instead of the usual error-based cost function). In this paper, we propose a Kalman filter with individual control factors, in terms of using a different level of uncertainty for each coefficient of the filter. As compared to the basic Kalman filter (which imposes the same uncertainty for all the coefficients of the impulse response), the proposed algorithm achieves better performance, especially in terms of the steady-state misalignment. Constantin Paleologu, Jacob Benesty, Silviu Ciochina, Steven L. Grant |
ICASSP | 2 |
| 2014 | On the noisereduction performance of the MVDR beamformer innoisy and reverberant environmentsabstractThe minimum variance distortionless response (MVDR) beam-former has been widely studied for extraction of desired speech signals in noisy acoustic environments. The performance of this beam-former, however, depends on many factors such as the array geometry, the source incidence angle, the noise field characteristics, the reverberation conditions, etc. In this paper, we study the performance of the MVDR beamformer in different noise and reverberation conditions with a linear microphone array. Using the gain in signal-to-noise ratio (SNR) as the performance metric, we show that the optimal performance of the MVDR beamformer generally occurs when the source is in the endfire directions in different types of noise, which indicates that, as long as a linear array is used, we should configure it in such a way that the endfire direction is pointed to the desired source. Simulations in reverberant environments also verified this result, though the performance difference between end-fire and broadside directions reduces as the degree of reverberation increases. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2014 | Examples of optimal noise reduction filters derived from the squared Pearson correlation coefficientabstractThis paper studies the problem of single-channel noise reduction in the time domain. Based on some orthogonal decomposition developed recently and the squared Pearson correlation coefficient (SPCC), several noise reduction filters are derived. We will show that the optimization of the SPCC leads to the Wiener, minimum variance distortionless response (MVDR), minimum noise (MN), minimum uncorrelated speech and noise (MUSN), and linearly constrained minimum variance (LCMV) filters. We also compare the Wiener and MVDR filters derived from the SPCC to their counterparts derived from the mean-square error (MSE) criterion. Simulations are provided to illustrate the performance of all the deduced noise reduction filters. Jiaolong Yu, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 2 |
| 2014 | Widely linear general Kalman filter for stereophonic acoustic echo cancellation
Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
Signal Process. | 2 |
| 2014 | A family of maximum SNR filters for noise reductionabstractThis paper is devoted to the study and analysis of the maximum signal-to-noise ratio (SNR) filters for noise reduction both in the time and short-time Fourier transform (STFT) domains with one single microphone and multiple microphones. In the time domain, we show that the maximum SNR filters can significantly increase the SNR but at the expense of tremendous speech distortion. As a consequence, the speech quality improvement, measured by the perceptual evaluation of speech quality (PESQ) algorithm, is marginal if any, regardless of the number of microphones used. In the STFT domain, the maximum SNR filters are formulated by considering the interframe information in every frequency band. It is found that these filters not only improve the SNR, but also improve the speech quality significantly. As the number of input channels increases so is the gain in SNR as well as the speech quality. This demonstrates that the maximum SNR filters, particularly the multichannel ones, in the STFT domain may be of great practical value. Gongping Huang, Jacob Benesty, Tao Long 0004, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Multichannel Noise Reduction in the Karhunen-Loève Expansion DomainabstractThe noise reduction problem is traditionally approached in the time, frequency, or transform domain. Having a signal dependent transform has shown some advantages over the traditional signal independent transform. Recently, the single-channel noise reduction problem in the Karhunen-Loève expansion (KLE) domain has received special attention. In this paper, the noise reduction problem in the KLE domain is studied from a multichannel perspective. We present a new formulation of the problem, in which inter-channel and inter-mode correlations are optimally exploited. We derive different optimal noise reduction filters and present a set of useful performance measures within this framework. The performance of the different filters is then evaluated through experiments in which not only noise but also competing speech sources are present. It is shown that the proposed multichannel formulation is more robust to competing speech sources than the single-channel approach and that a better compromise between noise reduction and speech distortion can be obtained. Yesenia Lacouture-Parodi, Emanuël A. P. Habets, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Performance Study of the MVDR Beamformer as a Function of the Source Incidence AngleabstractLinear microphone arrays combined with the minimum variance distortionless response (MVDR) beamformer have been widely studied in various applications to acquire desired signals and reduce the unwanted noise. Most of the existing array systems assume that the desired sources are in the broadside direction. In this paper, we study and analyze the performance of the MVDR beamformer as a function of the source incidence angle. Using the signal-to-noise ratio (SNR) and beampattern as the criteria, we investigate its performance in four different scenarios: spatially white noise, diffuse noise, diffuse-plus-white noise, and point-source-plus-white noise. The results demonstrate that the optimal performance of the MVDR beamformer occurs when the source is in the endfire directions for diffuse noise and point-source noise while its SNR gain does not depend on the signal incidence angle in spatially white noise. This indicates that most current systems may not fully exploit the potential of the MVDR beamformer. This analysis does not only help us better understand this algorithm, but also helps us design better array systems for practical applications. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Design of Robust Differential Microphone ArraysabstractDifferential microphone arrays (DMAs), due to their small size and enhanced directivity, are quite promising in speech enhancement applications. However, it is well known that differential beamformers have the drawback of white noise amplification, which is a major issue in the processing of wideband signals such as speech. In this paper, we focus on the design of robust DMAs. Based on the Maclaurin’s series approximation and frequency-independent beampatterns, the robust first-, second-, and third-order DMAs are proposed by using more microphones than the order plus one, and the corresponding minimum-norm filters are derived. Compared to the traditional DMAs, the proposed designs are more robust with respect to white noise amplification while they are capable of achieving similar directional gains. Liheng Zhao, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Multichannel acoustic echo suppressionabstractAcoustic echo suppression (AES) provides an attractive alternative to acoustic echo cancellation (AEC) techniques for full-duplex communication in low-complexity systems. However, so far AES techniques are commonly known to introduce significant distortions to the desired signal. Moreover, most traditional echo control techniques typically require accurately detecting the contribution of the near-end speaker to the microphone signal (“double talk”). The extension of AES techniques to the multichannel case usually assumes a symmetric system design which is often not fulfilled by typical scenarios. In this paper we propose a novel approach to multichannel acoustic echo suppression, which aims at extracting the near-end signal using a constraint for a distortionless output, without requiring a double-talk detector, or a symmetric system design. In addition to the above mentioned properties, the multichannel AES is also shown to overcome the known challenges in conventional multichannel acoustic echo control setups. Karim Helwani, Herbert Buchner, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2013 | A study of the MVDR filter for acoustic echo suppressionabstractThis paper studies an echo suppression approach to reducing the undesired echoes that result from the acoustic coupling between a loudspeaker and a microphone in duplex voice communication. The approach consists of four basic steps. First, both the loudspeaker and microphone signals are partitioned into small overlapping frames. Second, each frame is transformed into the short-time Fourier transform (STFT) domain. Third, a minimum variance distortionless response (MVDR) filter is designed in each subband by explicitly using the interframe signal correlation. This MVDR filter is then used to estimate the echo signal and the obtained estimate is subsequently subtracted from the microphone signal. Finally, the time-domain processed signal is constructed using the overlap-add technique with the inverse STFT. Experiments are performed and the results demonstrate that this proposed method can achieve significant amount of echo suppression in practical room environments. Jacob Benesty, Jingdong Chen, Karim Helwani, Herbert Buchner |
ICASSP | 2 |
| 2013 | Multichannel signal enhancement using non-causal, time-domain filtersabstractIn the vast amount of time-domain filtering methods for speech enhancement, the filters are designed to be causal. Recently, however, it was shown that the noise reduction and signal distortion capabilities of such single-channel filters can be improved by allowing the filters to be non-causal. While non-causal filters require knowledge of the future, they can be implemented in practice by introducing a short delay. In this paper, we generalize the idea of exploiting non-causality in optimal filter designs to the multichannel scenario. More specifically, a set of optimal, non-causal, multichannel filters for enhancement based on an orthogonal decomposition is proposed. The evaluation shows that there is a potential gain in noise reduction and signal distortion by introducing non-causality. Moreover, experiments on real-life speech show that we can improve the perceptual quality. Jesper Rindom Jensen, Mads Græsbøll Christensen, Jacob Benesty |
ICASSP | 3 |
| 2013 | Variationally diagonalized multichannel state-space frequency-domain adaptive filtering for acoustic echo cancellationabstractIn this contribution, we present a novel low-complexity state-space algorithm for multichannel acoustic echo cancellation. The reduction in complexity is brought about by means of top-down imposition of mutual independence on the respective acoustic echo paths within a variational Bayesian framework. This results in a fully diagonalized multichannel echo-path state estimator with a complexity that varies linearly with the channel order. The state estimator is augmented with learning rules for the model parameters that are optimal in the maximum-likelihood sense. We substantiate the efficacy of our formulation by means of simulation results in the presence of changes in the echo paths and continuous double-talk. Sarmad Malik, Jacob Benesty |
ICASSP | 2 |
| 2013 | Study of the optimal and simplified Kalman filters for echo cancellationabstractIn this paper, we study the time-domain Kalman filter in the context of echo cancellation. We explain the fundamental differences between the Kalman filter and the recursive least-squares (RLS) algorithm. Also, we show that the normalized least-mean-square (NLMS) algorithm has a clear relationship with the Kalman filter. Furthermore, a simplified Kalman filter is derived and by a judicious choice of its parameters, this algorithm behaves like a variable step-size adaptive filter. Simulation results indicate the good performance of the optimal and simplified Kalman filtering algorithms. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
ICASSP | 2 |
| 2013 | A widely linear model for stereophonic acoustic echo cancellation
Cristian Lucian Stanciu, Jacob Benesty, Constantin Paleologu, Tomas Gänsler, Silviu Ciochina |
Signal Process. | 2 |
| 2013 | A Single-Channel MVDR Filter for Acoustic Echo SuppressionabstractAcoustic echo suppression techniques for full-duplex communication in low-complexity systems are commonly known to introduce distortion to the desired signal (i.e., near-end speech). Moreover, most traditional echo control techniques typically require accurately detecting the contribution of the near-end speaker to the microphone signal (“double talk”). In this letter, we propose a novel approach to acoustic echo suppression, which aims at extracting the near-end signal using a constraint for minimizing the distortion, and without requiring a double-talk detector. Karim Helwani, Herbert Buchner, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2013 | On the Time-Domain Widely Linear LCMV Filter for Noise Reduction With a Stereo SystemabstractThis paper deals with the problem of noise reduction in stereo sound systems where the objective is not only to reduce noise, but also to preserve the spatial information of both the desired speech and noise sources so that the listener can still localize the speech and noise sources by listening to the enhanced binaural outputs. To achieve this objective, we use the widely linear (WL) framework developed previously and convert the problem of binaural noise reduction into one of monaural filtering with complex signals. We then present a way to decompose both the complex speech and noise signal vectors into two orthogonal components: one correlated and the other uncorrelated with the corresponding current signal sample. With this decomposition, the problem of noise reduction with preservation of the spatial information of speech and noise sources is formulated as an optimization problem with two constraints: one on the desired speech and the other on the preservation of the noise signal. We then derive a WL linearly constrained minimum variance (LCMV) filter, which can take advantage of the statistics and noncircularity of the complex speech signal to achieve noise reduction. In contrast to the WL Wiener and minimum variance distortionless response (MVDR) filters developed previously that can only preserve the characteristics and spatial information of the desired sound source, this new WL LCMV filter has the potential to reduce noise while preserving the characteristics and spatial information of both the desired and noise sources at the same time. Experimental results are provided to justify the claimed merits of the proposed WL LCMV filter. Jingdong Chen, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | A Two-Stage Beamforming Approach for Noise Reduction and DereverberationabstractIn general, the signal-to-noise ratio as well as the signal-to-reverberation ratio of speech received by a microphone decrease when the distance between the talker and microphone increases. Dereverberation and noise reduction algorithm are essential for many applications such as videoconferencing, hearing aids, and automatic speech recognition to improve the quality and intelligibility of the received desired speech that is corrupted by reverberation and noise. In the last decade, researchers have aimed at estimating the reverberant desired speech signal as received by one of the microphones. Although this approach has let to practical noise reduction algorithms, the spatial diversity of the received desired signal is not exploited to dereverberate the speech signal. In this paper, a two-stage beamforming approach is presented for dereverberation and noise reduction. In the first stage, a signal-independent beamformer is used to generate a reference signal which contains a dereverberated version of the desired speech signal as received at the microphones and residual noise. In the second stage, the filtered microphone signals and the noisy reference signal are used to obtain an estimate of the dereverberated desired speech signal. In this stage, different signal-dependent beamformers can be used depending on the desired operating point in terms of noise reduction and speech distortion. The presented performance evaluation demonstrates the effectiveness of the proposed two-stage approach. Emanuël A. P. Habets, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Multi-Microphone Noise Reduction Based on Orthogonal Noise Signal DecompositionsabstractMulti-microphone noise reduction plays an increasing and important role in acoustic communication systems. Existing multichannel noise reduction filters are commonly computed based on a single noise covariance matrix. Recently, an orthogonal noise signal decomposition was proposed that uses a single noise signal as a reference. Using this decomposition, it was possible to reformulate the noise reduction problem and derived a multichannel noise reduction filter that allows a tradeoff between the noise that is coherent and incoherent with respect to the reference signal. In this contribution, we analyze the previously proposed decomposition and propose an orthogonal decomposition that is based on a rank-one projection of all noise signals. The projection is chosen such that the total variance of the coherent noise component is maximized. To further improve the separation between coherent noise and incoherent noise, a rank-Qprojection of the observed noise signals is proposed. The decomposed noise covariance matrix is then used to derive a minimum variance distortionless response beamformer that allows a tradeoff between coherent and incoherent noise reduction, and to form a constraint matrix for a linearly constrained minimum variance beamformer. The results of the performance evaluation demonstrate the advantage of the proposed decompositions over the previously proposed decomposition. Emanuël A. P. Habets, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | A Class of Optimal Rectangular Filtering Matrices for Single-Channel Signal Enhancement in the Time DomainabstractIn this paper, we introduce a new class of optimal rectangular filtering matrices for single-channel speech enhancement. The new class of filters exploits the fact that the dimension of the signal subspace is lower than that of the full space. By doing this, extra degrees of freedom in the filters, that are otherwise reserved for preserving the signal subspace, can be used for achieving an improved output signal-to-noise ratio (SNR). Moreover, the filters allow for explicit control of the tradeoff between noise reduction and speech distortion via the chosen rank of the signal subspace. An interesting aspect is that the framework in which the filters are derived unifies the ideas of optimal filtering and subspace methods. A number of different optimal filter designs are derived in this framework, and the properties and performance of these are studied using both synthetic, periodic signals and real signals. The results show a number of interesting things. Firstly, they show how speech distortion can be traded for noise reduction and vice versa in a seamless manner. Moreover, the introduced filter designs are capable of achieving both the upper and lower bounds for the output SNR via the choice of a single parameter. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Study of the General Kalman Filter for Echo CancellationabstractThe Kalman filter is a very interesting signal processing tool, which is widely used in many practical applications. In this paper, we study the Kalman filter in the context of echo cancellation. The contribution of this work is threefold. First, we derive a different form of the Kalman filter by considering, at each iteration, a block of time samples instead of one time sample as it is the case in the conventional approach. Second, we show how this general Kalman filter (GKF) is connected with some of the most popular adaptive filters for echo cancellation, i.e., the normalized least-mean-square (NLMS) algorithm, the affine projection algorithm (APA) and its proportionate version (PAPA). Third, a simplified Kalman filter is developed in order to reduce the computational load of the GKF; this algorithm behaves like a variable step-size adaptive filter. Simulation results indicate the good performance of the proposed algorithms, which can be attractive choices for echo cancellation. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A multichannel widely linear approach to binaural noise reduction using an array of microphonesabstractThis paper deals with the problem of binaural noise reduction using an array of microphones. This is a very important problem in applications such as teleconferencing and hearing aids where there is a need to mitigate the noise effect from the noisy signals picked up by multiple microphones and produce two “clean” outputs. The mitigation of the noise should be made in such a way that no audible distortion is added to the two outputs (this is the same as in the single-channel case) and meanwhile the spatial information of the desired sound source should be preserved so that, after noise reduction, the listener will still be able to localize the sound source thanks to his/her binaural hearing mechanism. In this paper, we present a novel approach to this problem where we first form a number of complex input signals from the multiple and real microphone observations. We also merge the two expected real outputs into a complex output signal. The widely linear estimation theory is then used to derive optimal noise reduction filters that can achieve noise reduction while preserving the desired signal (speech) and its spatial information. With this new formulation, the Wiener and minimum variance distortionless response (MVDR) filters are derived. Experiments are provided to justify the effectiveness of these filters. Jacob Benesty, Jingdong Chen |
ICASSP | 1 |
| 2012 | Single-channel noise reduction in the STFT domain based on the bifrequency spectrumabstractThis paper studies the problem of noise reduction in the short-time Fourier transform (STFT) domain. Traditionally, the STFT coefficients in different frequency bands are assumed to be independent. This assumption holds when the signals are stationary and the fast Fourier transform(FFT) length is sufficiently large. In practice, however, speech is nonstationary and also the FFT length cannot be very large due to practical reasons. So, there always exists some correlation between STFT coefficients from neighboring frequency bands. An important question then arises: how the interband correlation can be used to optimize noise reduction performance? This paper addresses this issue. We discuss two solutions in the framework of the bifrequency spectrum. One considers the cross-correlation between all the frequency bands and the other takes into account only the cross-correlation between neighboring bands. While the former is optimal from a theoretical perspective, the latter is more practical as it is more immune to the error in correlation matrix estimation. Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2012 | Multi-microphone noise reduction using interchannel and interframe correlationsabstractMulti-microphone noise reduction methods often operate in the time-frequency domain in which a complex gain is applied to each time-frame and subband. These methods can achieve good noise reduction with little speech distortion by exploiting the fact that the desired signal is correlated across the channels. In the context of single-microphone noise reduction, it has been shown recently that the performance in terms of noise reduction and speech distortion can be improved by exploiting the correlation between subsequent time-frames, i.e., by exploiting the interframe correlation. In this paper, we exploit both interchannel and interframe correlations in the context of multi-microphone noise reduction. Now the interframe correlation is taken into account, i.e., a filter is applied in each subband and channel instead of just a gain. The results of our experimental study show that we can improve the fullband signal-to-noise ratios (SNRs) by using interchannel and interframe correlations when dealing with signals, such as speech, that exhibit a sufficiently large interframe correlation. Emanuël A. P. Habets, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2012 | Multichannel noise reduction wiener filter in the Karhunen-Loève expansion domainabstractThis paper explores the noise reduction problem in the Karhunen-Loève expansion (KLE) domain from a multichannel perspective. Based on formulations proposed for the design of optimal single-channel noise reduction in the KLE domain, we formulate the multichannel noise reduction in the KLE domain. Two different performance measures are presented: the noise reduction and speech distortion. The optimal multichannel Wiener filter is derived and its performance in terms of noise reduction and speech distortion is compared with the performance of the optimal single-channel Wiener filter. Experimental results show that a significant improvement in performance is obtained when using multiple microphone signals. The multichannel Wiener filter results also in better noise reduction in the presence of coherent noise sources. Yesenia Lacouture-Parodi, Emanuël A. P. Habets, Jacob Benesty |
ICASSP | 3 |
| 2012 | Optimal rectangular filtering matrix for noise reduction in the time domainabstractIn this paper, we study the noise reduction problem in the time domain and present a frame-based method to decompose the clean speech vector into two orthogonal components: one correlated and the other uncorrelated with the current desired speech vector to be estimated. In comparison with the sample-based decomposition developed in the previous research that uses only forward prediction, this new decomposition exploits both the forward prediction and interpolation. Based on this new decomposition, we formulate different optimization cost functions and address the issue of how to design Wiener and minimum variance distortionless response (MVDR) filtering matrices by optimizing these new cost functions. We also discuss the relationship between the Wiener and MVDR filtering matrices and show that the MVDR filtering matrix can achieve noise reduction without adding speech distortion; but it reduces less noise than the Wiener filtering matrix. Compared with the sample-based algorithms developed in the previous study, the proposed frame-based algorithms can achieve better noise reduction performance. Furthermore, they are computationally more efficient, and therefore, more suitable for practical implementation. Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2012 | Regularization of the improved proportionate affine projection algorithmabstractIn sparse adaptive filters, the adaptation gain is “proportionately” redistributed among all the coefficients, emphasizing the large ones in order to speed up their convergence. The improved proportionate affine projection algorithm (IPAPA) is a very attractive choice for echo cancellation, since it combines the good convergence features of the affine projection algorithm (APA) and the gain factors of the improved proportionate normalized least-mean-square (IPNLMS) algorithm. Similar to the APA, a matrix inversion is required within the IPAPA. For practical reasons, the matrix needs to be regularized before inversion, i.e., a positive constant is added to the elements of its main diagonal. In this paper, we propose a formula for choosing the regularization parameter of the IPAPA, aiming at attenuating the effects of the noise in the adaptive filter estimate. Simulation results indicate the validity of this approach in both network and acoustic echo cancellation scenarios. Constantin Paleologu, Jacob Benesty, Felix Albu |
ICASSP | 2 |
| 2012 | On an iterative method for basis pursuit with application to echo cancellation with sparse impulse responsesabstractBasis pursuit has been shown to be an effective method of solving inverse problems with a small amount of data when the system to be determined has a sparse representation. Adaptive filters fall under this general category of problems. Here, we use the echo cancellation context to introduce a method of solving the basis pursuit problem with an iterative method based on the proportionate normalized affine projection algorithm (PAPA). Earlier, it has been shown that PAPA can be derived from a basis pursuit perspective. Here we refine the assumptions made in those derivations and show that an iterative form of PAPA yields the same results as basis pursuit without resorting to the simplex method. The resulting algorithm has extremely fast convergence for adaptive filters with very sparse impulse responses. Simulations using the new iterative approach are also presented. Steven L. Grant, Jacob Benesty |
ICASSP | 3 |
| 2012 | A novel perspective on stereophonic acoustic echo cancellationabstractThe stereophonic acoustic echo is due to the coupling between two loudspeakers and two microphones. In the classical approach, this configuration is modelled by a two-input/two-output system with real random variables. In this paper, we propose to redesign this scheme as a single-input/single-output system with complex random variables. In this framework, we illustrate the behavior of some basic adaptive algorithms and present a distortion method which is more suitable for this model. Cristian Lucian Stanciu, Jacob Benesty, Constantin Paleologu, Tomas Gänsler, Silviu Ciochina |
ICASSP | 2 |
| 2012 | Proportionate affine projection algorithms from a basis pursuit perspectiveabstractIt can be shown that the update of the affine projection algorithm (APA) can be decomposed as the sum of two orthogonal vectors. One of these vectors is derived from an ℓ2-norm optimization problem while the other one is simply a good initialization vector. By replacing the ℓ2-norm optimization with the ℓk-norm optimization (with 0k-norm optimizations over the performance of the PAPAs. Constantin Paleologu, Jacob Benesty |
ISCAS | 2 |
| 2012 | A Perspective on Differential Microphone Arrays in the Context of Noise ReductionabstractIn this correspondence, we study the performance of differential microphone arrays (DMAs) in terms of noise reduction, speech distortion, and signal-to-noise ratio (SNR) gain. We also investigate their beampatterns and array gains. We start by establishing the expressions of these performance measures involving general derivatives of the channel transfer functions. Afterwards, we specify our results in the case of anechoic near-field and far-field propagation models. Jacob Benesty, Mehrez Souden, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | A Perspective on Frequency-Domain Beamformers in Room AcousticsabstractSignals captured by a set of microphones in a speech communication system are mixtures of desired signals and noise. In this paper, a different perspective on frequency-domain beamformers in room acoustics is provided. Specifically, the observed noise signals are divided into coherent and incoherent signal components while no assumptions are being made regarding the number of coherent noise sources and the noise sound field. From this perspective, performance measures are defined and existing beamformers are deduced. In addition, a new and general tradeoff beamformer is proposed that enables a compromise between noise reduction and speech distortion on the one hand, and coherent noise versus incoherent noise reductions on the other hand. The presented performance evaluation shows how existing beamformers and the tradeoff beamformer perform in different scenarios. Emanuël A. P. Habets, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A Speech Distortion and Interference Rejection Constraint BeamformerabstractSignals captured by a set of microphones in a speech communication system are mixtures of desired and undesired signals and ambient noise. Existing beamformers can be divided into those that preserve or distort the desired signal. Beamformers that preserve the desired signal are, for example, the linearly constrained minimum variance (LCMV) beamformer that is supposed, ideally, to reject the undesired signal and reduce the ambient noise power, and the minimum variance distortionless response (MVDR) beamformer that reduces the interference-plus-noise power. The multichannel Wiener filter, on the other hand, reduces the interference-plus-noise power without preserving the desired signal. In this paper, a speech distortion and interference rejection constraint (SDIRC) beamformer is derived that minimizes the ambient noise power subject to specific constraints that allow a tradeoff between speech distortion and interference-plus-noise reduction on the one hand, and undesired signal and ambient noise reductions on the other hand. Closed-form expressions for the performance measures of the SDIRC beamformer are derived and the relations to the aforementioned beamformers are derived. The performance evaluation demonstrates the tradeoffs that can be made using the SDIRC beamformer. Emanuël A. P. Habets, Jacob Benesty, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A Multi-Frame Approach to the Frequency-Domain Single-Channel Noise Reduction ProblemabstractThis paper focuses on the class of single-channel noise reduction methods that are performed in the frequency domain via the short-time Fourier transform (STFT). The simplicity and relative effectiveness of this class of approaches make them the dominant choice in practical systems. Over the past years, many popular algorithms have been proposed. These algorithms, no matter how they are developed, have one feature in common: the solution is eventually formulated as a gain function applied to the STFT of the noisy signal only in the current frame, implying that the interframe correlation is ignored. This assumption is not accurate for speech enhancement since speech is a highly self-correlated signal. In this paper, by taking the interframe correlation into account, a new linear model for speech spectral estimation and some optimal filters are proposed. They include the multi-frame Wiener and minimum variance distortionless response (MVDR) filters. With these filters, both the narrowband and fullband signal-to-noise ratios (SNRs) can be improved. Furthermore, with the MVDR filter, speech distortion at the output can be zero. Simulations present promising results in support of the claimed merits obtained by theoretical analysis. Yiteng Huang, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Non-Causal Time-Domain Filters for Single-Channel Noise ReductionabstractIn many existing time-domain filtering methods for noise reduction in, e.g., speech processing, the filters are causal. Such causal filters can be implemented directly in practice. However, it is possible to improve the performance of such noise reduction filtering methods in terms of both noise suppression and signal distortion by allowing the filters to be non-causal. Non-causal time-domain filters require knowledge of the future, and are therefore not directly implementable. If the observed signal is processed in blocks, however, the non-causal filters are implementable. In this paper, we propose such non-causal time-domain filters for noise reduction in speech applications. We also propose some performance measures that enable us to evaluate the performance of non-causal filters. Moreover, it is shown how some of the filters can be updated recursively. Using the recursive expressions, it is also shown that the output SNRs of the filters always increase as we increase the length of the filter when the desired signal is stationary. From both the theoretical and practical evaluations of the filters, it is clearly shown that the performance of time-domain filtering methods for noise reduction can be improved by introducing non-causality. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Enhancement of Single-Channel Periodic Signals in the Time-DomainabstractMost state-of-the-art filtering methods for speech enhancement require an estimate of the noise statistics, but the noise statistics are difficult to estimate in practice when speech is present. Thus, nonstationary noise will have a detrimental impact on the performance of most speech enhancement filters. The impact of such noise can be reduced by using the signal statistics rather than the noise statistics in the filter design. For example, this is possible by assuming a harmonic model for the desired signal; while this model fits well for voiced speech, it will not be appropriate for unvoiced speech. That is, signal-dependent methods based on the signal statistics will introduce undesired distortion for some parts of speech compared to signal-independent methods based on the noise statistics. Since both the signal-independent and signal-dependent approaches to speech enhancement have advantages, it is relevant to combine them to reduce the impact of their individual disadvantages. In this paper, we give theoretical insights into the relationship between these different approaches, and these reveal a close relationship between the two approaches. This justifies joint use of such filtering methods which can be beneficial from a practical point of view. Our experimental results confirm that both signal-independent and signal-dependent approaches have advantages and that they are closely-related. Moreover, as a part of our experiments, we illustrate the practical usefulness of combining signal-independent and signal-dependent enhancement methods by applying such methods jointly on real-life speech. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A variable step size evolutionary affine projection algorithmabstractIt is well known that the affine projection algorithm (APA) offers a good tradeoff between convergence rate/tracking and computational complexity. Recently, the evolutionary APA (E-APA) with a variable projection order has been proposed. In this paper, we propose a variable step size (VSS) version of the E-APA, called VSS-E-APA. It is shown that the VSS-E-APA is robust to near-end signal variations. Also, it has both a fast convergence speed and a small steady-state error and a much reduced numerical complexity than the VSS-APA. Felix Albu, Constantin Paleologu, Jacob Benesty |
ICASSP | 3 |
| 2011 | A single-channel noise reduction MVDR filterabstractMost existing approaches for single-channel noise reduction in the frequency domain via the short-time Fourier transform (STFT) assume that consecutive time-frames are uncorrelated with each other. As a result, algorithms based on this assumption do not, obviously, take the interframe correlation into account. In this paper, we pro pose a framework that considers this interframe correlation. An important consequence in including this useful information, is that now it is possible to derive a (single-channel) minimum variance distortionless response (MVDR) filter for noise reduction. The experimental study shows impressive results. Jacob Benesty, Yiteng Huang |
ICASSP | 1 |
| 2011 | On single-channel noise reduction in the time domainabstractIn this paper, we revisit the noise-reduction problem in the time domain and present a way to decompose the filtered speech into two uncorrelated (orthogonal) components: the desired speech and the interference. Based on this new decomposition, we discuss how to form different optimization cost functions and address the issue of how to design different noise-reduction filters by optimizing these new cost functions. Particularly, we cover the design of the maximum signal-to-noise-ratio (SNR), the Wiener, the minimum variance distortionless response (MVDR), and the tradeoff filters. It is interesting that with this new decomposition, we can now design the MVDR filter that can achieve noise reduction without adding speech distortion in the single-channel case, which has never been seen before. We also demonstrate that the maximum SNR, Wiener, and tradeoff filters are identical to the MVDR filter up to a scaling factor. From a theoretical point of view, this scaling factor is not significant and should not affect the output SNR at any processing time. But from a practical viewpoint, the scaling factor can be time-varying due to the nonstationarity of the speech and possibly the noise and can cause discontinuity in the residual noise level, which is unpleasant to listen to. As a result, it is essential to have the scaling factor right from one processing sample (or frame) to another in order to avoid large distortions and for this reason, it is recommended to use the MVDR filter in speech enhancement applications. Jingdong Chen, Jacob Benesty, Yiteng Huang, Tomas Gänsler |
ICASSP | 2 |
| 2011 | An efficient variable step-size proportionate affine projection algorithmabstractProportionate-type affine projection algorithms (PAPAs) are very attractive choices for echo cancellation. These algorithms combine the good features (convergence and tracking) of the affine projection algorithm (APA) and the proportionate idea, which exploits the sparseness character of the echo path in order to further increase their convergence rate. In this paper, we develop a variable step-size (VSS) version of a recently proposed PAPA. The new algorithm achieves a good compromise between fast convergence rate and low misadjustment, but also has a low computational complexity as compared to its classical counterparts. Constantin Paleologu, Jacob Benesty, Felix Albu, Silviu Ciochina |
ICASSP | 2 |
| 2011 | Class of double-talk detectors based on the holder inequalityabstractMost of the echo cancellers are equipped with a double-talk detector (DTD) in order to control the behavior of the adaptive filter during double-talk situations. In this paper, we propose a class of DTDs based on the Holder inequality. These DTDs are simple to implement, have low computational complexity, and perform well even for low echo-to-noise ratios. As a particular case, it is shown that the well-known Geigel algorithm can be obtained from this approach. Constantin Paleologu, Jacob Benesty, Tomas Gänsler, Silviu Ciochina |
ICASSP | 2 |
| 2011 | Stability analysis of adaptive filters with regression vector nonlinearities
Leonardo Rey Vega, Hernan Rey, Jacob Benesty |
Signal Process. | 3 |
| 2011 | Binaural Noise Reduction in the Time Domain With a Stereo SetupabstractBinaural noise reduction with a stereophonic (or simply stereo) setup has become a very important problem as stereo sound systems and devices are being more and more deployed in modern voice communications. This problem is very challenging since it requires not only the reduction of the noise at the stereo inputs, but also the preservation of the spatial information embodied in the two channels so that after noise reduction the listener can still localize the sound source from the binaural outputs. As a result, simply applying a traditional single-channel noise reduction technique to each channel individually may not work as the spatial effects may be destroyed. In this paper, we present a new formulation of the binaural noise reduction problem in stereo systems. We first form a complex signal from the stereo inputs with one channel being its real part and the other being its imaginary part. By doing so, the binaural noise reduction problem can be processed by a single-channel widely linear filter. The widely linear estimation theory is then used to derive optimal noise reduction filters that can fully take advantage of the noncircularity of the complex speech signal to achieve noise reduction while preserving the desired signal (speech) and spatial information. With this new formulation, the Wiener, minimum variance distortionless response (MVDR), maximum signal-to-noise ratio (SNR), and tradeoff filters are derived. Experiments are provided to justify the effectiveness of these filters. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | On Regularization in Adaptive FilteringabstractRegularization plays a fundamental role in adaptive filtering. An adaptive filter that is not properly regularized will perform very poorly. In spite of this, regularization in our opinion is underestimated and rarely discussed in the literature of adaptive filtering. There are, very likely, many different ways to regularize an adaptive filter. In this paper, we propose one possible way to do it based on a condition that intuitively makes sense. From this condition, we show how to regularize four important algorithms: the normalized least-mean-square (NLMS), the signed-regressor NLMS (SR-NLMS), the improved proportionate NLMS (IPNLMS), and the SR-IPNLMS. Jacob Benesty, Constantin Paleologu, Silviu Ciochina |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | An Integrated Solution for Online Multichannel Noise Tracking and ReductionabstractNoise statistics estimation is a paramount issue in the design of reliable noise-reduction algorithms. Although significant efforts have been devoted to this problem in the literature, most developed methods so far have focused on the single-channel case. When multiple microphones are used, it is important that the data from all the sensors are optimally combined to achieve judicious updates of the noise statistics and the noise-reduction filter. This contribution is devoted to the development of a practical approach to multichannel noise tracking and reduction. We combine the multichannel speech presence probability (MC-SPP) that we proposed in an earlier contribution with an alternative formulation of the minima-controlled recursive averaging (MCRA) technique that we generalize from the single-channel to the multichannel case. To demonstrate the effectiveness of the proposed MC-SPP and multichannel noise estimator, we integrate them into three variants of the multichannel noise reduction Wiener filter. Experimental results show the advantages of the proposed solution. Mehrez Souden, Jingdong Chen, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Study of the widely linear Wiener filter for noise reductionabstractThis paper develops a new widely linear noise-reduction Wiener filter based on the variance and pseudo-variance of the short-time Fourier transform coefficients of speech signals. We show that this new noise-reduction filter has many interesting properties, including but not limited to: 1) it causes less speech distortion as compared to the classical noise-reduction Wiener filter; 2) its minimum mean-squared error (MSE) is smaller than that of the classical Wiener filter; 3) it can increase the subband signal-to-noise ratio (SNR), while the classical Wiener filter has no effect on the subband SNR for any given signal frame and subband. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP | 1 |
| 2010 | Analysis of the frequency-domain Wiener filter with the prediction gainabstractThis paper presents a theoretical analysis on the performance of the optimal noise-reduction filter in the frequency domain. Using the autoregressive (AR) model to model both the clean speech and noise, we build the relationship between the Wiener filter and the AR parameters of the clean speech and noise signals. We show that if noise is not predictable, the Wiener filter is mostly related to the AR parameters of the desired speech signal. On the contrary, if the desired signal is not predictable, the Wiener filter is then mostly related to the AR parameters of the noise signal. More importantly, we provide the bounds for noise reduction, speech distortion, and SNR improvement, and show that the performance of the Wiener filter in terms of SNR improvement and degree of noise reduction and speech distortion is closely related to the prediction gain of the desired speech and noise signals. Jingdong Chen, Jacob Benesty, Yiteng Huang |
ICASSP | 2 |
| 2010 | An improved proportionate NLMS algorithm based on the l0 normabstractThe proportionate normalized least-mean-square (PNLMS) algorithm was developed in the context of network echo cancellation. It has been proven to be efficient when the echo path is sparse, which is not always the case in real-world echo cancellation. The improved PNLMS (IPNLMS) algorithm is less sensitive to the sparseness character of the echo path. This algorithm uses the l1norm to exploit sparseness of the impulse response that needs to be identified. In this paper, we propose an IPNLMS algorithm based on the l0norm, which represents a better measure of sparseness than the l1norm. Simulation results prove that the proposed algorithm outperforms the original IPNLMS algorithm. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
ICASSP | 2 |
| 2010 | A variable step-size normalized sign algorithm for acoustic echo cancelationabstractA variable step size normalized sign algorithm (VSS-NSA) is proposed, for acoustic echo cancelation, which adjusts its step size automatically by matching the L1norm of the a posteriori error to that of the background noise plus near-end signal. Simulation results show that the new algorithm combined with double-talk detection outperforms the dual sign algorithm (DSA) and the normalized triple-state sign algorithm (NTSSA) in terms of convergence rate and stability. Tiange Shao, Yahong Rosa Zheng, Jacob Benesty |
ICASSP | 3 |
| 2010 | Eigenanalysis-based broadband source localizationabstractThe localization process consists in finding the candidate source location that maximizes the synchrony between the properly time-shifted microphone outputs. In addition to using well-known crosscorrelation-based criteria such as the steered response power (SRP), minimum variance (MV), and multichannel crosscorrelation (MCCC), this synchrony can be measured using the averaged magnitude difference function (AMDF) and the averaged magnitude sum function (AMSF) whose calculations involve low computational cost. In this paper, we study the crosscorrelation and AMDF (with AMSF) based approaches using an arbitrary number of microphones. Specifically, we use the eigenanalysis of the parameterized spatial correlation matrix (PSCM) to first provide a unifying study of the most popular crosscorrelation-based techniques. Then, we show the efficiency of the AMDF and AMSF in localizing an acoustic source using multiple microphones by proposing two new parameterized matrices named as the parameterized averaged magnitude difference matrix (PAMDM) and the parameterized averaged magnitude sum matrix (PAMSM). The eigenanalysis of these two matrices reveals new criteria. Mehrez Souden, Jacob Benesty, Sofiène Affes |
ICASSP | 2 |
| 2010 | Linear filtering for noise reduction and interference rejectionabstractWe study the linearly constrained minimum variance (LCMV) and the minimum variance distortionless response (MVDR) filters when multiple interferers and unknown (ambient) noise coexist with a target speech signal. Precisely, the LCMV is designed to remove all the interference signals while preserving the desired speech and attempting to reduce the ambient noise components. The MVDR is simply formulated such that the overall ambient-noise-plus-interference are reduced while satisfying a distortionless constraint. We provide simplified expressions for both beamformers and show their relationship. Furthermore, we underline the limitations of the LCMV when the ambient noise is present. When the latter is absent, we also prove that the MVDR degenerates to the LCMV. Numerical examples are provided to support our study. Mehrez Souden, Jacob Benesty, Sofiène Affes |
ICASSP | 2 |
| 2010 | A robust variable step-size affine projection algorithm
Leonardo Rey Vega, Hernan Rey, Jacob Benesty |
Signal Process. | 3 |
| 2010 | On widely linear Wiener and tradeoff filters for noise reduction
Jacob Benesty, Jingdong Chen, Yiteng Huang |
Speech Commun. | 1 |
| 2010 | A Widely Linear Distortionless Filter for Single-Channel Noise ReductionabstractTraditionally in the single-channel noise-reduction problem, speech distortion is inevitable since the desired signal is also filtered while filtering the noise. In fact, the more the noise is reduced, the more the speech distortion is added into the desired signal, as proved in the literature. So, if we require no speech distortion, we either end up with no noise reduction at all or have to use multiple sensors. In this paper, we attempt to apply the widely linear (WL) estimation theory to noise reduction. Unlike the traditional approaches that only filter the short-time Fourier transform (STFT) of the noisy signal, the method developed in this paper applies the noise-reduction filter to both the STFT of the noisy signal and its conjugate. With the constraint of no speech distortion, a WL distortionless filter is derived. We show that this new optimal filter can fully take advantage of the noncircularity property of speech signals to achieve up to 3-dB signal-to-noise-ratio (SNR) improvement without introducing any speech distortion, which can only be obtained with the traditional approaches if two or more microphones are used. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Signal Process. Lett. | 1 |
| 2010 | Proportionate Adaptive Filters From a Basis Pursuit PerspectiveabstractIn this letter, we show that the normalized least-mean-square (NLMS) algorithm and the affine projection algorithm (APA) can be decomposed as the sum of two orthogonal vectors. One of these vectors is derived from an ℓ2-norm optimization problem while the other one is simply a good initialization vector. By replacing this optimization with the basis pursuit, which is based on the ℓ1-norm optimization, we derive the proportionate NLMS (PNLMS) algorithm and the proportionate APA (PAPA). Many other adaptive filters can be derived following this approach, including new ones. Jacob Benesty, Constantin Paleologu, Silviu Ciochina |
IEEE Signal Process. Lett. | 1 |
| 2010 | An Efficient Proportionate Affine Projection Algorithm for Echo CancellationabstractProportionate-type normalized least-mean-square algorithms were developed in the context of echo cancellation. In order to further increase the convergence rate and tracking, the “proportionate” idea was applied to the affine projection algorithm (APA) in a straightforward manner. The objective of this letter is twofold. First, a general framework for the derivation of proportionate-type APAs is proposed. Second, based on this approach, a new proportionate-type APA is developed, taking into account the “history” of the proportionate factors. The benefit is also twofold. Simulation results indicate that the proposed algorithm outperforms the classical one (achieving faster tracking and lower misadjustment). Besides, it also has a lower computational complexity due to a recursive implementation of the “proportionate history.” Constantin Paleologu, Silviu Ciochina, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2010 | An Affine Projection Sign Algorithm Robust Against Impulsive InterferencesabstractA new affine projection sign algorithm (APSA) is proposed, which is robust against non-Gaussian impulsive interferences and has fast convergence. The conventional affine projection algorithm (APA) converges fast at a high cost in terms of computational complexity and it also suffers performance degradation in the presence of impulsive interferences. The family of sign algorithms (SAs) stands out due to its low complexity and robustness against impulsive noise. The proposed APSA combines the benefits of the APA and SA by updating its weight vector according to theL1-norm optimization criterion while using multiple projections. The features of the APA and theL1-norm minimization guarantee the APSA an excellent candidate for combatting impulsive interference and speeding up the convergence rate for colored inputs at a low computational complexity. Simulations in a system identification context show that the proposed APSA outperforms the normalized least-mean-square (NLMS) algorithm, APA, and normalized sign algorithm (NSA) in terms of convergence rate and steady-state error. The robustness of the APSA against impulsive interference is also demonstrated. Tiange Shao, Yahong Rosa Zheng, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2010 | On the Global Output SNR of the Parameterized Frequency-Domain Multichannel Noise Reduction Wiener FilterabstractThe parameterized multichannel Wiener filter (PMWF) is known to allow for a flexible tuning of speech distortion and noise reduction. In addition, the output signal-to-noise ratio (SNR) is a natural metric that shows the true effect of such a filter on both residual noise and filtered speech. Earlier contributions have only shown that the output SNR of the Wiener filter is larger than the input SNR (as a proof of its effectiveness). However, the effect of the tuning parameter (denoted as ß following the notation of [1]) on this metric is not yet understood. In this paper, we prove that the global (fullband) output SNR of the PMWF is an increasing function of ß but remains below an asymptotic value. As a byproduct, a very simplified proof of the global output SNR improvement is provided. Mehrez Souden, Jacob Benesty, Sofiène Affes |
IEEE Signal Process. Lett. | 2 |
| 2010 | New Insights Into the MVDR Beamformer in Room AcousticsabstractThe minimum variance distortionless response (MVDR) beamformer, also known as Capon's beamformer, is widely studied in the area of speech enhancement. The MVDR beamformer can be used for both speech dereverberation and noise reduction. This paper provides new insights into the MVDR beamformer. Specifically, the local and global behavior of the MVDR beamformer is analyzed and novel forms of the MVDR filter are derived and discussed. In earlier works it was observed that there is a tradeoff between the amount of speech dereverberation and noise reduction when the MVDR beamformer is used. Here, the tradeoff between speech dereverberation and noise reduction is analyzed thoroughly. The local and global behavior, as well as the tradeoff, is analyzed for different noise fields such as, for example, a mixture of coherent and non-coherent noise fields, entirely non-coherent noise fields and diffuse noise fields. It is shown that maximum noise reduction is achieved when the MVDR beamformer is used for noise reduction only. The amount of noise reduction that is sacrificed when complete dereverberation is required depends on the direct-to-reverberation ratio of the acoustic impulse response between the source and the reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction. When desiring both speech dereverberation and noise reduction, the results also demonstrate that the amount of noise reduction that is sacrificed decreases when the number of microphones increases. Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot, Jacek Dmochowski |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | On Optimal Frequency-Domain Multichannel Linear Filtering for Noise ReductionabstractSeveral contributions have been made so far to develop optimalmultichannellinear filtering approaches and show their ability to reduce the acoustic noise. However, there has not been a clearunifying theoreticalanalysis of their performance in terms of both noise reduction and speech distortion. To fill this gap, we analyze the frequency-domain (non-causal) multichannel linear filtering for noise reduction in this paper. For completeness, we consider the noise reduction constrained optimization problem that leads to the parameterized multichannel non-causal Wiener filter (PMWF). Our contribution is fivefold. First, we formally show that the minimum variance distortionless response (MVDR) filter is a particular case of the PMWF by properly formulating the constrained optimization problem of noise reduction. Second, we propose new simplified expressions for the PMWF, the MVDR, and the generalized sidelobe canceller (GSC) that depend on the signals' statistics only. In contrast to earlier works, these expressions are explicitly independent of the channel transfer function ratios. Third, we quantify the theoretical gains and losses in terms of speech distortion and noise reduction when using the PWMF by establishing new simplified closed-form expressions for three performance measures, namely, the signal distortion index, the noise reduction factor (originally proposed in the paper titled ldquoNew insights into the noise reduction Wiener filter,rdquo by J. Chen (IEEE Transactions on Audio, Speech, and Language Processing, Vol. 15, no. 4, pp. 1218-1234, Jul. 2006) to analyze the single channel time-domain Wiener filter), and the output signal-to-noise ratio (SNR). Fourth, we analyze the effects of coherent and incoherent noise in addition to the benefits of utilizing multiple microphones. Fifth, we propose a new proof for thea posterioriSNR improvement achieved by the PMWF. Finally, we provide some simulations results to corroborate the findings of this work. Mehrez Souden, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Broadband Source Localization From an Eigenanalysis PerspectiveabstractBroadband source localization has several applications ranging from automatic video camera steering to target signal tracking and enhancement through beamforming. Consequently, there has been a considerable amount of effort to develop reliable methods for accurate localization over the last few decades. Essentially, the localization process consists in finding the candidate source location that maximizes the synchrony between the properly time-shifted microphone outputs. In addition to using well known cross-correlation-based criteria such as the steered response power (SRP), minimum variance (MV), and multichannel cross-correlation (MCCC), this synchrony can also be measured using the averaged magnitude difference function (AMDF) and the averaged magnitude sum function (AMSF) whose calculations involve low computational cost. In earlier related works, the latter techniques have been used for time delay estimation (TDE) of a target source observed by only one pair of microphones. Their generalization to the multiple microphone case and application to source localization have not been studied yet. In this paper, we consider both categories, i.e., cross-correlation and AMDF (with AMSF)-based approaches, using an arbitrary number of microphones, and analyze their performance. Specifically, we first provide a unifying study of the most popular cross-correlation-based techniques, such as the SRP, MV, and MCCC. In this paper, we use the eigenanalysis of the parameterized spatial correlation matrix (PSCM) to classify these methods and gain some insight into their performance. We demonstrate, for instance, that the MV and SRP consist in searching the major eigenvalue of the PSCM, while the MCCC, essentially, combines its minor eigenvalues when scanning for the source location. Inspired by this analysis, we show, in the second part of this work, the efficiency of the AMDF and AMSF in localizing an acoustic source using multiple microphones. Indeed, we propose two new parameterized matrices named as the parameterized averaged magnitude difference matrix (PAMDM) and the parameterized averaged magnitude sum matrix (PAMSM). The eigenanalysis of these matrices also reveals new criteria for acoustic source localization. Simulation results are provided to illustrate the effectiveness of all the investigated and proposed methods. Mehrez Souden, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Gaussian Model-Based Multichannel Speech Presence ProbabilityabstractThe knowledge of the target speech presence probability in a mixture of signals captured by a speech communication system is of paramount importance in several applications including reliable noise reduction algorithms. In this correspondence, we establish a new expression for speech presence probability when an array of microphones with an arbitrary geometry is used. Our study is based on the assumption of the Gaussian statistical model for all signals and involves the noise and noisy data statistics only. In comparison with the single-channel case, the new proposed multichannel approach can significantly increase the detection accuracy. In particular, when the additive noise is spatially coherent, perfect speech presence detection is theoretically possible, while when the noise is spatially white, a coherent summation of speech components is performed to allow for enhanced speech presence probability estimation. Mehrez Souden, Jingdong Chen, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | On noise reduction in the Karhunen-Loève expansion domainabstractIn this paper, we study the noise-reduction problem in the Karhunen-Loève expansion domain. We develop two classes of optimal filters. The first class estimates a frame of speech by filtering the corresponding frame of the noisy speech. We will show that several well-known existing methods belong or are closely related to this category. The second class, which has not been studied before, obtains noise reduction by filtering not only the current frame, but also a number of previous consecutive frames of the noisy speech. We will discuss how to design the optimal noise-reduction filters in each class and demonstrate the properties of the deduced optimal filters. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP | 1 |
| 2009 | On a tradeoff between dereverberation and noise reduction using the MVDR beamformerabstractThe minimum variance distortionless response (MVDR) beamformer can be used for both speech dereverberation and noise reduction. In this paper we analyse the tradeoff between the amount of speech dereverberation and noise reduction achieved by the MVDR beamformer. We show that the amount of noise reduction that is sacrificed when desiring both speech dereverberation and noise reduction depends on the direct-to-reverberation ratio of the acoustic transfer function between the desired source and a reference microphone. The performance evaluation supports the theoretical analysis and demonstrates the tradeoff between speech dereverberation and noise reduction. Emanuël A. P. Habets, Jacob Benesty, Israel Cohen, Sharon Gannot |
ICASSP | 2 |
| 2009 | Using the Pearson correlation coefficient to develop an optimally weighted cross relation based blind SIMO identification algorithmabstractBlind SIMO identification is challenging when additive noise is strong and for ill-conditioned/acoustic SIMO systems. A weighted cross relation (CR) algorithm presumably can be robust to noise but there lacks a practical way to define the weights. In this paper, the Pearson correlation coefficient (PCC) is used to develop an optimally weighted CR algorithm, which is validated by simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2009 | New insights into non-causal multichannel linear filtering for noise reductionabstractWe investigate a general framework for noise reduction which consists in controlling the level of signal distortion while reducing the level of noise. A parameterized non-causal filter that allows for tuning the signal distortion and noise reduction inversely is obtained and is referred to as parameterized multichannel non-causal Wiener filter (PMWF) herein. The same optimization problem leads to the minimum variance distortionless response (MVDR) as a particular case of the PMWF. In contrast to earlier works, the proposed expressions of the PMWF and MVDR are simplified and require the knowledge of the speech and noise statistics only. To rigorously quantify the gains and losses when using these filters, we establish simplified closed-form expressions for three measures, namely, the signal distortion index, the noise reduction factor, and the output signal-to-noise ratio (SNR), and highlight the tradeoff between noise reduction and speech distortion in the multichannel case. Mehrez Souden, Jacob Benesty, Sofiène Affes |
ICASSP | 2 |
| 2009 | Study of the quaternion LMS and four-channel LMS algorithmsabstractThe recently proposed quaternion least-mean-square (QLMS) algorithm for adaptive filtering of three- and four-dimensional signals has been analysed in the context of multi-step ahead prediction. For rigour, the relationship between multichannel LMS (MLMS) and QLMS is examined, and their differences are highlighted. This is achieved both in terms of the input-output relationship and in terms of the dynamics of weight updates. The convergence of QLMS is investigated and stability bounds confirm that QLMS and MLMS are fundamentally different. Simulations on both synthetic and real world multidimensional signals support the analysis. Clive Cheong Took, Danilo P. Mandic, Jacob Benesty |
ICASSP | 3 |
| 2009 | Noise Reduction Algorithms in a Generalized Transform DomainabstractNoise reduction for speech applications is often formulated as a digital filtering problem, where the clean speech estimate is obtained by passing the noisy speech through a linear filter/transform. With such a formulation, the core issue of noise reduction becomes how to design an optimal filter (based on the statistics of the speech and noise signals) that can significantly suppress noise without introducing perceptually noticeable speech distortion. The optimal filters can be designed either in the time or in a transform domain. The advantage of working in a transform space is that, if the transform is selected properly, the speech and noise signals may be better separated in that space, thereby enabling better filter estimation and noise reduction performance. Although many different transforms exist, most efforts in the field of noise reduction have been focused only on the Fourier and Karhunen-Loeve transforms. Even with these two, no formal study has been carried out to investigate which transform can outperform the other. In this paper, we reformulate the noise reduction problem into a more generalized transform domain. We will show some of the advantages of working in this generalized domain, such as 1) different transforms can be used to replace each other without any requirement to change the algorithm (optimal filter) formulation, and 2) it is easier to fairly compare different transforms for their noise reduction performance. We will also address how to design different optimal and suboptimal filters in such a generalized transform domain. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Study of the Noise-Reduction Problem in the Karhunen-LoÈve Expansion DomainabstractNoise reduction, which aims at estimating a clean speech from a noisy observation, has long been an active research area. The standard approach to this problem is to obtain the clean speech estimate by linearly filtering the noisy signal. The core issue, then, becomes how to design an optimal linear filter that can significantly suppress noise without introducing perceptually noticeable speech distortion. Traditionally, the optimal noise-reduction filters are formulated in either the time or the frequency domains. This paper studies the problem in the Karhunen–LoÈve expansion domain. We develop two classes of optimal filters. The first class achieves a frame of speech estimate by filtering the corresponding frame of the noisy speech. We will show that many existing methods such as the widely used Wiener filter and subspace technique are closely related to this category. The second class obtains noise reduction by filtering not only the current frame, but also a number of previous consecutive frames of the noisy speech. We will discuss how to design the optimal noise-reduction filters in each class and demonstrate, through both theoretical analysis and experiments, the properties of the deduced optimal filters. Jingdong Chen, Jacob Benesty, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | An Information-Theoretic Viewof ArrayProcessingabstractThe removal of noise and interference from an array of received signals is a most fundamental problem in signal processing research. To date, many well-known solutions based on second-order statistics (SOS) have been proposed. This paper views the signal enhancement problem as one of maximizing the mutual information between the source signal and array output. It is shown that if the signal and noise are Gaussian, the maximum mutual information estimation (MMIE) solution is not unique but consists of an infinite set of solutions which encompass the SOS-based optimal filters. The application of the MMIE principle to Laplacian signals is then examined by considering the important problem of estimating a speech signal from a set of noisy observations. It is revealed that while speech (well modeled by a Laplacian distribution) possesses higher order statistics (HOS), the well-known SOS-based optimal filters maximize the Laplacian mutual information as well; that is, the Laplacian mutual information differs from the Gaussian mutual information by a single term whose dependence on the beamforming weights is negligible. Simulation results verify these findings. Jacek Dmochowski, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | A Family of Robust Algorithms Exploiting Sparsity in Adaptive FiltersabstractWe introduce a new family of algorithms to exploit sparsity in adaptive filters. It is based on a recently introduced new framework for designing robust adaptive filters. It results from minimizing a certain cost function subject to a time-dependent constraint on the norm of the filter update. Although in general this problem does not have a closed-form solution, we propose an approximate one which is very close to the optimal solution. We take a particular algorithm from this family and provide some theoretical results regarding the asymptotic behavior of the algorithm. Finally, we test it in different environments for system identification and acoustic echo cancellation applications. Leonardo Rey Vega, Hernan Rey, Jacob Benesty, Sara Tressens |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | A minimum speech distortion multichannel algorithm for noise reductionabstractNoise reduction using multiple microphones remains a challenging and crucial research problem. This paper presents a new multichannel noise-reduction algorithm based on spatio-temporal prediction. Unlike many multichannel techniques that attempt to achieve both speech dereverberation and noise reduction at the same time, this new approach puts aside speech dereverberation and formulates the problem as one of estimating the speech component received at one microphone using the observations from all the available microphones. In comparison with the existing techniques such as beamforming, this new multichannel approach has many appealing properties: it does not require the knowledge of the source location or the channel impulse responses; the multiple microphones do not have to be arranged into a specific array geometry; it works the same for both the far-field and near-field cases; and most importantly, it can produce very good noise reduction with minimum speech distortion in real acoustic environments. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP | 1 |
| 2008 | Fast steered response power source localization using inverse mapping of relative delaysabstractThe acoustic source localization problem has significance for modern intelligent communication systems and future human-computer interaction applications. Although steered-beamforming localization methods perform well in adverse conditions, the algorithms are inherently computationally burdensome due to their need to search the entire location space. This paper presents a computationally- reduced version of the steered response power (SRP) method. The proposed approach is rooted in the inverse mapping of relative delays to candidate source locations, which allows for the transformation of the iterative search from the multidimensional location space to the one-dimensional relative delay space. By subsetting the set of traversed relative delays to only those that experience a high-level of cross-correlation, the computational load is reduced by up to 90 % without incurring a loss in localization accuracy. Jacek Dmochowski, Jacob Benesty, Sofiène Affes |
ICASSP | 2 |
| 2008 | Generalized crosstalk cancellation and equalization using multiple loudspeakers for 3D sound reproduction at the ears of multiple listenersabstractA classical crosstalk cancellation and equalization (CTCE) system uses two loudspeakers and assumes only one listener. In this paper, the idea is generalized to the design of using multiple loudspeakers for multiple listeners. We will show that if the number of loudspeakers is equal to the number of ears, only a least-squares (LS) solution can be obtained, while using more loudspeakers than ears we have more options: either an LS solution or an exact solution for perfect CTCE. Via simulations using real impulse responses measured in the varechoic chamber at Bell Labs, we learn that a CTCE system employing more loudspeakers will be more robust to errors in the estimated acoustic impulse responses. Yiteng Huang, Jacob Benesty, Jingdong Clien |
ICASSP | 2 |
| 2008 | Double-talk robust VSS-NLMS algorithm for under-modeling acoustic echo cancellationabstractMost of the adaptive algorithms used for acoustic echo cancellation (AEC) are designed assuming an exact modeling scenario (i.e., the acoustic echo path and the adaptive filter have the same length) and a single-talk context (i.e., the near-end speech is absent). In real-world AEC applications, the adaptive filter works most likely in an under-modeling situation, i.e., its length is smaller than the length of the acoustic impulse response, so that the under-modeling noise is present. Also, the double-talk case is almost inherent, so that a double-talk detector (DTD) is usually involved. Both aspects influence and limit the algorithm’s performance. Taking into account these two practical issues, a double-talk robust variable step size normalized least-meansquare (VSS-NLMS) algorithm is proposed in this paper. This algorithm is nonparametric in the sense that it does not require any information about the acoustic environment, so that it is robust and easy to control in practice. Constantin Paleologu, Silviu Ciochina, Jacob Benesty |
ICASSP | 3 |
| 2008 | Microphone arrays for noise reduction with low signal distortion in room acousticsabstractWe propose a new method for noise reduction using a microphone array. The method takes advantage of the spatial diversity inherent to microphone arrays and optimizes certain criteria, namely, the output signal to noise ratio (SNR) or the mean squared error (MSE), subject to the constraint of spatial prediction that relates the noise free signals captured by the microphones. Simulation results demonstrate that resorting to this new method leads to high rate of noise reduction and low signal distortion. Mehrez Souden, Jacob Benesty, Sofiène Affes |
ICASSP | 2 |
| 2008 | Linear System Identification Based on a Third-Order Tensor DecompositionabstractA wide variety of system identification problems can be efficiently addressed based on the Kronecker product decomposition of the impulse response, together with low-rank approximations. Such an approach solves the original system identification problem using a combination of two shorter filters. In this paper, targeting a higher dimensionality reduction, we develop a solution based on a third-order tensor decomposition. In addition, the problem of approximating the rank of a tensor is avoided thanks to the control of a matrix rank. Then, an iterative Wiener filter is developed, which outperforms both the conventional benchmark and the previously developed counterpart that exploits the second-order decomposition. Jacob Benesty, Constantin Paleologu, Silviu Ciochina |
IEEE Signal Process. Lett. | 1 |
| 2008 | Design of Steerable Linear Differential Microphone Arrays With Omnidirectional and Bidirectional SensorsabstractThis paper is dedicated to the design of fully steerable linear differential microphone arrays (LDMAs). We analyze the steerable ideal spatial responses and explain why conventional LDMAs consisting of only omnidirectional microphones have limited steering ability. In order to circumvent this limitation, we suggest to use both omnidirectional and bidirectional (with a dipole shaped directivity pattern) microphones. We discuss the minimum numbers of omnidirectional and bidirectional sensors required for achieving steerable spatial responses and present a method to design fully steerable differential beamformers with LDMAs through the Jacobi-Anger series expansion. Simulations validate the presented technique and the steering flexibility of the designed LDMAs. Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 5 |
| 2008 | A Robust Variable Forgetting Factor Recursive Least-Squares Algorithm for System IdentificationabstractThe performance of the recursive least-squares (RLS) algorithm is governed by the forgetting factor. This parameter leads to a compromise between (1) the tracking capabilities and (2) the misadjustment and stability. In this letter, a variable forgetting factor RLS (VFF-RLS) algorithm is proposed for system identification. In general, the output of the unknown system is corrupted by a noise-like signal. This signal should be recovered in the error signal of the adaptive filter after this one converges to the true solution. This condition is used to control the value of the forgetting factor. The simulation results indicate the good performance and the robustness of the proposed algorithm. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
IEEE Signal Process. Lett. | 2 |
| 2008 | Variable Step-Size NLMS Algorithm for Under-Modeling Acoustic Echo CancellationabstractIn acoustic echo cancellation (AEC) applications, where the acoustic echo paths are extremely long, the adaptive filter works most likely in an under-modeling situation. Most of the adaptive algorithms for AEC were derived assuming an exact modeling scenario, so that they do not take into account the under-modeling noise. In this letter, a variable step-size normalized least-mean-square (VSS-NLMS) algorithm suitable for the under-modeling case is proposed. This algorithm does not require any a priori information about the acoustic environment; as a result, it is very robust and easy to control in practice. The simulation results indicate the good performance of the proposed algorithm. Constantin Paleologu, Silviu Ciochina, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2008 | On the Importance of the Pearson Correlation Coefficient in Noise ReductionabstractNoise reduction, which aims at estimating a clean speech from noisy observations, has attracted a considerable amount of research and engineering attention over the past few decades. In the single-channel scenario, an estimate of the clean speech can be obtained by passing the noisy signal picked up by the microphone through a linear filter/transformation. The core issue, then, is how to find an optimal filter/transformation such that, after the filtering process, the signal-to-noise ratio (SNR) is improved but the desired speech signal is not noticeably distorted. Most of the existing optimal filters (such as the Wiener filter and subspace transformation) are formulated from the mean-square error (MSE) criterion. However, with the MSE formulation, many desired properties of the optimal noise-reduction filters such as the SNR behavior cannot be seen. In this paper, we present a new criterion based on the Pearson correlation coefficient (PCC). We show that in the context of noise reduction the squared PCC (SPCC) has many appealing properties and can be used as an optimization cost function to derive many optimal and suboptimal noise-reduction filters. The clear advantage of using the SPCC over the MSE is that the noise-reduction performance (in terms of the SNR improvement and speech distortion) of the resulting optimal filters can be easily analyzed. This shows that, as far as noise reduction is concerned, the SPCC-based cost function serves as a more natural criterion to optimize as compared to the MSE. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | A Minimum Distortion Noise Reduction Algorithm With Multiple MicrophonesabstractThe problem of noise reduction using multiple microphones has long been an active area of research. Over the past few decades, most efforts have been devoted to beamforming techniques, which aim at recovering the desired source signal from the outputs of an array of microphones. In order to work reasonably well in reverberant environments, this approach often requires such knowledge as the direction of arrival (DOA) or even the room impulse responses, which are difficult to acquire reliably in practice. In addition, beamforming has to compromise its noise reduction performance in order to achieve speech dereverberation at the same time. This paper presents a new multichannel algorithm for noise reduction, which formulates the problem as one of estimating the speech component observed at one microphone using the observations from all the available microphones. This new approach explicitly uses the idea of spatial–temporal prediction and achieves noise reduction in two steps. The first step is to determine a set of inter-sensor optimal spatial–temporal prediction transformations. These transformations are then exploited in the second step to form an optimal noise-reduction filter. In comparison with traditional beamforming techniques, this new method has many appealing properties: it does not require DOA information or any knowledge of either the reverberation condition or the channel impulse responses; the multiple microphones do not have to be arranged into a specific array geometry; it works the same for both the far-field and near-field cases; and, most importantly, it can produce very good and robust noise reduction with minimum speech distortion in practical environments. Furthermore, with this new approach, it is possible to apply postprocessing filtering for additional noise reduction when a specified level of speech distortion is allowed. Jingdong Chen, Jacob Benesty, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Linearly Constrained Minimum Variance Source Localization and Spectral EstimationabstractA signal's spectrum is a representation of the signal in terms of elementary basis functions which facilitates the extraction of desired information. For a temporal signal, the spectrum is one-dimensional and expresses the time-domain signal as a linear combination of sinusoidal basis functions. A space-time signal possesses a multidimensional Fourier transform known as the wavenumber-frequency spectrum, which represents the space-time signal as a weighted summation of monochromatic plane waves. The spatial and temporal frequencies are not separable, as spatial frequency is itself a function of the temporal frequency. Thus, it seems natural to analyze and estimate the spatial and temporal frequency components in tandem. It is therefore surprising that conventional spectral estimation methods focus on either the spatial or temporal dimension, without any regard for the other. Spatial spectral estimation is commonly referred to as source localization, as the direction of the wavenumber vector is indeed the direction of propagation. Conventional methods analyze a solely spatial aperture without accounting for the temporal structure of the desired signal. Conversely, temporal spectral estimation is performed using a single sensor, and thus the signal aperture is purely temporal. This paper proposes a spatiotemporal framework for spectral estimation based on the linearly constrained minimum variance (LCMV) beamforming method proposed by Frost in 1972. The aperture consists of an array of sensors, each storing a set of previous temporal samples. It is first shown that by taking into account the temporal structure of the desired signal, the ensuing source location estimate is more robust to the effects of noise and reverberation. Unlike conventional localizers, the LCMV steered beamtemporallyfocuses the array onto the desired signal. The desired signal is modeled by an autoregressive (AR) process, and the resulting AR coefficients are embedded in the linear constraints. As a result, the rate of anomalous estimates is significantly reduced as compared to existing techniques. Moreover, it is then demonstrated that by employing multiple sensors and steering the array to the assumed source location, the estimate of the desired signal's temporal spectrum contains a lesser contribution from the unwanted noise and reverberation. Jacek Dmochowski, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Analysis and Comparison of Multichannel Noise Reduction Methods in a Common FrameworkabstractNoise reduction for speech enhancement is a useful technique, but in general it is a challenging problem. While a single-channel algorithm is easy to use in practice, it inevitably introduces speech distortion to the desired speech signal while reducing noise. Today, the explosive growth in computational power and the continuous drop in the cost and size of acoustic electric transducers are driving the interest of employing multiple microphones in speech processing systems. This opens new opportunities for noise reduction. In this paper, we present an analysis of three multichannel noise reduction algorithms, namely Wiener filter, subspace, and spatial-temporal prediction, in a common framework. We intend to investigate whether it is possible for the multichannel noise reduction algorithms to reduce noise without speech distortion. Finally, we justify what we learn via theoretical analyses by simulations using real impulse responses measured in the varechoic chamber at Bell Labs. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | A Variable Step-Size Affine Projection Algorithm Designed for Acoustic Echo CancellationabstractThe adaptive algorithms used for acoustic echo cancellation (AEC) have to provide (1) high convergence rates and good tracking capabilities, since the acoustic environments imply very long and time-variant echo paths, and (2) low misadjustment and robustness against background noise variations and double-talk. In this context, the affine projection algorithm (APA) and different versions of it are very attractive choices for AEC. However, an APA with a constant step-size parameter has to compromise between the performance criteria (1) and (2). Therefore, a variable step-size APA (VSS-APA) represents a more reliable solution. In this paper, we propose a VSS-APA derived in the context of AEC. Most of the APAs aim to cancelp(i.e., projection order) previous a posteriori errors at every step of the algorithm. The proposed VSS-APA aims to recover the near-end signal within the error signal of the adaptive filter. Consequently, it is robust against near-end signal variations (including double-talk). This algorithm does not require anyaprioriinformation about the acoustic environment, so that it is easy to control in practice. The simulation results indicate the good performance of the proposed algorithm as compared to other members of the APA family. Constantin Paleologu, Jacob Benesty, Silviu Ciochina |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | On Recursive and Fast Recursive Computation of the Capon SpectrumabstractThe Capon spectrum, which is known to have better resolution than the periodogram, has been widely used in various applications. Normally, the Capon spectrum is estimated through the direct computation of the inverse of the data correlation (or covariance) matrix. This so-called direct inverse approach is, however, computationally very expensive due to the high computational cost involved in the matrix inversion. This paper deals with fast and efficient algorithms in computing the Capon spectrum. Inspired from the recursive idea established in the area of adaptive signal processing, we first derive a recursive Capon algorithm. This new algorithm does not require an explicit matrix inversion, and hence is more efficient to implement than the direct inverse method. We then develop a fast version of the recursive algorithm, which can further reduce the complexity of the recursive one by an order of magnitude. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP (3) | 1 |
| 2007 | An Acoustic MIMO Framework for Analyzing Microphone-Array BeamformingabstractAlthough a significant amount of research attention has been devoted to microphone-array beamforming, the performance of all the developed algorithms in practical acoustic environments is still far from meeting our expectation. So further research efforts on this topic are indispensable. In this paper, we treat a microphone array as a multiple-input multiple-output (MIMO) system and develop a general framework for analyzing performance of beamforming algorithms based on the acoustic MIMO channel impulse responses. Under this framework, we study the bounds for the length of beamforming filter, which in turn shows the performance bounds of beamforming in terms of speech dereverberation and interference suppression. We also discuss the intrinsic relationships among different classical beamforming techniques and explain, from the channel condition point of view, what the prerequisites have to be fulfilled in order for those techniques to work. Jingdong Chen, Jacob Benesty, Yiteng Huang |
ICASSP (1) | 2 |
| 2007 | Direction of Arrival Estimation using Eigenanalysis of the Parameterized Spatial Correlation MatrixabstractThe estimation of the direction-of-arrival (DOA) of one or more acoustic sources is an area that has generated much interest in recent years, with applications like automatic video camera steering and multi-party stereophonic teleconferencing entering the market. Time-difference-of-arrival (TDOA) based methods compute each relative delay using only two microphones, even though additional microphones are usually available, and thus suffer from the effects of background noise and reverberation. This paper deals with DOA estimation based on spatial spectral estimation, and proposes a novel DOA estimator based on the eigenvalues of the parameterized spatial correlation matrix. Simulation results confirm the ability of the proposed method to provide reliable estimates even in heavily reverberant environments. Jacek Dmochowski, Jacob Benesty, Sofiène Affes |
ICASSP (1) | 2 |
| 2007 | Laplace Entropy and its Application to Time Delay Estimation for Speech SignalsabstractTime delay estimation (TDE) is a basic technique for numerous applications where there is a need to localize and track a radiating source. It is particularly challenging in the presence of noise and reverberation, and when the source signal is speech which is inherently nonstationary and random. The most important TDE algorithms for two sensors are based on the generalized cross-correlation (GCC) method. These algorithms perform reasonably well when reverberation or noise is not too high. In an earlier study of the authors, a more sophisticated approach was proposed. It employs more sensors and takes advantage of their delay redundancy to improve the precision of the TDOA (time difference of arrival) estimate between the first two sensors. The approach is based on the multichannel cross-correlation coefficient (MCCC) and was found more robust to noise and reverberation. In this paper, we show that this approach can also be developed on a basis of joint entropy. For Gaussian signals, we show that, in the search of the TDOA estimate, maximizing MCCC is equivalent to minimizing joint entropy. But with the generalization of the idea to non-Gaussian speech signals, the joint entropy based new multichannel TDE algorithm manifests a potential to outperform the MCCC-based method. Since there is no rigorous mathematical formula for speech entropy, we use the assumption that speech can be plausibly modeled by a Laplace distribution and develop a practical approximation of Laplace entropy for TDE of speech signals. The performance of the proposed new algorithm is investigated via simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (1) | 2 |
| 2007 | Estimating the Two-Dimensional Coherence FunctionabstractIn this paper, we extend the one-dimensional Capon-based magnitude square coherence (MSC) spectral estimator, to form two-dimensional Capon- and APES-based MSC spectral estimators. The resulting estimators are found to yield significantly improved estimates as compared to the typical Welch-based estimator. Furthermore, we introduce a computationally efficient time-updating of the presented MSC estimators, exploiting their inherent time-varying displacement structure. The presented updating is found to dramatically lower the computational requirement of reevaluating the MSC spectral estimates. Andreas Jakobsson, Stephen R. Alty, Jacob Benesty |
ICASSP (3) | 3 |
| 2007 | A Robust Adaptive Filtering Algorithm Against Impulsive NoiseabstractA new framework for designing robust adaptive filters is introduced. It is based on the optimization of a certain cost function subject to a time-dependent constraint on the norm of the filter update. Particularly, we will derive a robust variable step-size NLMS algorithm which optimizes the square norm of the a posteriori error subject to the constraint on the norm of the filter change. We also show the link between the proposed algorithm and another one derived using a robust statistics approach. The algorithm is then tested in different environments for system identification and acoustic echo cancelation applications. Leonardo Rey Vega, Hernan Rey, Jacob Benesty, Sara Tressens |
ICASSP (3) | 3 |
| 2007 | On the optimal linear filtering techniques for noise reduction
Jingdong Chen, Jacob Benesty, Yiteng Huang |
Speech Commun. | 2 |
| 2007 | Time Delay Estimation via Minimum EntropyabstractTime delay estimation (TDE) is a basic technique for numerous applications where there is a need to localize and track a radiating source. The most important TDE algorithms for two sensors are based on the generalized cross-correlation (GCC) method. These algorithms perform reasonably well when reverberation or noise is not too high. In an earlier study by the authors, a more sophisticated approach was proposed. It employs more sensors and takes advantage of their delay redundancy to improve the precision of the time difference of arrival (TDOA) estimate between the first two sensors. The approach is based on the multichannel cross-correlation coefficient (MCCC) and was found more robust to noise and reverberation. In this letter, we show that this approach can also be developed on a basis of joint entropy. For Gaussian signals, we show that, in the search of the TDOA estimate, maximizing MCCC is equivalent to minimizing joint entropy. However, with the generalization of the idea to non-Gaussian signals (e.g., speech), the joint entropy-based new TDE algorithm manifests a potential to outperform the MCCC-based method Jacob Benesty, Yiteng Huang, Jingdong Chen |
IEEE Signal Process. Lett. | 1 |
| 2007 | On Crosstalk Cancellation and Equalization With Multiple Loudspeakers for 3-D Sound ReproductionabstractPeople prefer to be able to enjoy spatial audio without wearing a headphone. Such a tethered device is anyway inconvenient and undesirable, if not cumbersome. Alternatively, 3D sound can be delivered to a listener with loudspeakers. However, crosstalk arises, and the rendered binaural signals are distorted by room reverberation when arriving at the listener's two ears, which lead to the need for a crosstalk cancellation and equalization (CTCE) system. Classical CTCE systems employ only two loudspeakers, and their performance is usually unsatisfactory in practice. While the idea of using more loudspeakers has been investigated, it was never shown why using more loudspeakers is theoretically more advantageous for CTCE. In this letter, we will study this problem and demonstrate that with two loudspeakers, only a least-squares (LS) solution can be obtained, while using multiple loudspeakers, we have more options: either an LS solution or an exact solution for perfect CTCE. These findings are justified by simulations using real impulse responses measured in the varechoic chamber at Bell Labs. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 2 |
| 2007 | On Microphone-Array Beamforming From a MIMO Acoustic Signal Processing PerspectiveabstractAlthough many microphone-array beamforming algorithms have been developed over the past few decades, most such algorithms so far can only offer limited performance in practical acoustic environments. The reason behind this has not been fully understood and further research on this matter is indispensable. In this paper, we treat a microphone array as a multiple-input multiple-output (MIMO) system and study its signal-enhancement performance. Our major contribution is fourfold. First, we develop a general framework for analyzing performance of beamforming algorithms based on the acoustic MIMO channel impulse responses. Second, we study the bounds for the length of the beamforming filter, which in turn shows the performance bounds of beamforming in terms of speech dereverberation and interference suppression. Third, we address the connection between beamforming and the multiple-input/output inverse theorem (MINT). Finally, we discuss the intrinsic relationships among different classical beamforming techniques and explain, from the channel condition perspective, what the prerequisites are for those techniques to work. Jacob Benesty, Jingdong Chen, Yiteng Huang, Jacek Dmochowski |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Direction of Arrival Estimation Using the Parameterized Spatial Correlation MatrixabstractThe estimation of the direction-of-arrival (DOA) of one or more acoustic sources is an area that has generated much interest in recent years, with applications like automatic video camera steering and multiparty stereophonic teleconferencing entering the market. DOA estimation algorithms are hindered by the effects of background noise and reverberation. Methods based on the time-differences-of-arrival (TDOA) are commonly used to determine the azimuth angle of arrival of an acoustic source. TDOA-based methods compute each relative delay using only two microphones, even though additional microphones are usually available. This paper deals with DOA estimation based on spatial spectral estimation, and establishes the parameterized spatial correlation matrix as the framework for this class of DOA estimators. This matrix jointly takes into account all pairs of microphones, and is at the heart of several broadband spatial spectral estimators, including steered-response power (SRP) algorithms. This paper reviews and evaluates these broadband spatial spectral estimators, comparing their performance to TDOA-based locators. In addition, an eigenanalysis of the parameterized spatial correlation matrix is performed and reveals that such analysis allows one to estimate the channel attenuation from factors such as uncalibrated microphones. This estimate generalizes the broadband minimum variance spatial spectral estimator to more general signal models. A DOA estimator based on the multichannel cross correlation coefficient (MCCC) is also proposed. The performance of all proposed algorithms is included in the evaluation. It is shown that adding extra microphones helps combat the effects of background noise and reverberation. Furthermore, the link between accurate spatial spectral estimation and corresponding DOA estimation is investigated. The application of the minimum variance and MCCC methods to the spatial spectral estimation problem leads to better resolution than that of the commonly used fixed-weighted SRP spectrum. However, this increased spatial spectral resolution does not always translate to more accurate DOA estimation. Jacek Dmochowski, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | A Generalized Steered Response Power Method for Computationally Viable Source LocalizationabstractThe process of locating an acoustic source given measurements of the sound field at multiple microphones is of significant interest as both a classical array signal processing problem, and more recently, as a solution to the problems of automatic camera steering, teleconferencing, hands-free processing, and others. Despite the proven efficacy of steered-beamformer approaches to localization in harsh conditions, their practical application to real-time settings is hindered by undesirably high computational demands. This paper presents a computationally viable implementation of the steered response power (SRP) source localization method. The conventional approach is generalized by introducing an inverse mapping that maps relative delays to sets of candidate locations. Instead of traversing the three-dimensional location space, the one-dimensional relative delay space is traversed; at each lag, all locations which are inverse mapped by that delay are updated. This means that the computation of the SRP map is no longer performed sequentially in space. Most importantly, by subsetting the space of relative delays to only those that achieve a high level of cross-correlation, the required number of algorithm updates is drastically reduced without compromising localization accuracy. The generalization is scalable in the sense that the level of subsetting is an algorithm parameter. It is shown that this generalization may be viewed as a spatial decomposition of the SRP energy map into weighted basis functions-in this context, it becomes evident that the full SRP search considers all basis functions (even the ones with very low weighting). On the other hand, it is shown that by only including a few basis functions per microphone pair, the SRP map is quite accurately represented. As a result, in a real environment, the proposed generalization achieves virtually the same anomaly rate as the full SRP search while only performing 10% the amount of algorithm updates as the full search. Jacek Dmochowski, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Enhancement of Spatial Sound Quality: A New Reverberation-Extraction Audio UpmixerabstractA system for the extraction of uncorrelated reverberation from two-channel (stereo) audio signals is proposed and evaluated. Applications for the new system vary from surround-sound multichannel loudspeaker upmixers for home-theater or automotive audio systems, to headphone-based auralization for enhancing the spatial sound quality of a listening experience. The new system uses the normalized-least-mean-square (NLMS) algorithm to equalize the two input signals with respect to both spectral magnitude and phase before differencing to remove correlated components. A theoretical model of the system based on a stochastic room impulse response model was validated by empirical measurements made in a reverberant hall with a microphone pair, and from a formal subjective evaluation the system is shown to be an effective approach to extracting reverberation from audio recordings. J. Usher, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Estimation of the Coherence Function with the MVDR ApproachabstractThe minimum variance distortionless response (MVDR), originally developed by Capon for frequency-wavenumber analysis, is a very well established method in array processing. It is also used in spectral estimation. The aim of this paper is to show how the MVDR method can be used to estimate the magnitude squared coherence (MSC) function, which is very useful in so many applications but so few methods exist to estimate it. Simulations show that our algorithm gives much more reliable results than the one based on the popular Welch's method. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP (3) | 1 |
| 2006 | Speech Acquisition and Enhancement in a Reverberant, Cocktail-Party-Like EnvironmentabstractDeveloping a successful multi-microphone speech acquisition system in a reverberant, cocktail-party-like environment is a very challenging problem since both interfering sources and reverberation need to be well controlled. In this paper, we propose an algorithm based on blind SIMO identification. We first blindly identify the channels from the interfering sources to all the microphones. Then we extract the speech signal of interest. Finally speech dereverberation is performed using the MINT method. Simulations with acoustic impulse responses measured in the varechoic chamber at Bell Labs are carried out to verify the proposed algorithm. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (5) | 2 |
| 2006 | Effect of Interchannel Coherence on Conditioning and Misalignment Performance for Stereo Acoustic ECHO CancellationabstractIt is well known that the performance in terms of misalignment of adaptive algorithms, in general, is dependent on the conditioning of the input signal covariance matrix. For two-channel (stereophonic) adaptive algorithms, this performance is further degraded by the high interchannel coherence between the two input signals. In this paper, we establish the relationship between interchannel coherence of the two input signals and condition of the corre- corresonding covariance matrix for stereo acoustic echo cancellation application. We further show how this relationship affects the misalignment performance of a two-channel frequency-domain adaptive algorithm. We provide simulation results for both WGN and speech input to verify our mathematical analysis. Andy W. H. Khong, Jacob Benesty, Patrick A. Naylor |
ICASSP (5) | 2 |
| 2006 | Optimum Variable Explicit Regularized Affine Projection AlgorithmabstractA variable regularized affine projection algorithm (VR-APA) is introduced, which does not require the classical step size. Its use is supported from different points of view. First, it has the property of being Hinfinoptimal, providing robust behavior against perturbations and model uncertainties. Second, the time varying regularization parameter is obtained by maximizing the speed of convergence of the algorithm. At each time step, it needs knowledge of the power of the estimation error vector, which can be estimated by averaging observable quantities. Although we first derive it for a linear time invariant (LTI) system, we show that the same expression holds if we consider a time varying system following a first order Markov model. Simulation results are presented to test the performance of the proposed algorithm and to compare it with other schemes under different situations Hernan Rey, Leonardo Rey Vega, Sara Tressens, Jacob Benesty |
ICASSP (3) | 4 |
| 2006 | A New Approach to Blind Separation of Two Sources with Three SensorsabstractAn exhaustive investigation of the overdetermined blind source separation (BSS) 3 times 2 problem is presented. We establish a new relationship between the column vectors of the channel matrix using second order statistics (SOS) only. This relationship is then exploited to factorize the separating matrix into a whitening term, whose expression depends on whether the observations are corrupted by noise or not, and a 2 times 2 orthogonal matrix that can be determined using high order statistics (HOS). Simulation results demonstrate that resorting to the new form of the whitening matrix, mainly in the case of noisy ill-conditioned BSS, results in improved performance compared to the common whitening-based techniques. Mehrez Souden, Sofiène Affes, Jacob Benesty |
VTC Fall | 3 |
| 2006 | The fast normalized cross-correlation double-talk detector
Tomas Gänsler, Jacob Benesty |
Signal Process. | 2 |
| 2006 | Identification of acoustic MIMO systems: Challenges and opportunities
Yiteng Huang, Jacob Benesty, Jingdong Chen |
Signal Process. | 2 |
| 2006 | A Nonparametric VSS NLMS AlgorithmabstractThe aim of a variable step size normalized least-mean-square (VSS-NLMS) algorithm is to try to solve the conflicting requirement of fast convergence and low misadjustment of the NLMS algorithm. Numerous VSS-NLMS algorithms can be found in the literature with a common point for most of them: they may not work very reliably since they depend on several parameters that are not simple to tune in practice. The objective of this letter is twofold. First, we explain a simple and elegant way to derive VSS-NLMS-type algorithms. Second, a new nonparametric VSS-NLMS is proposed that is easy to control and gives good performances in the context of acoustic echo cancellation Jacob Benesty, Hernan Rey, Leonardo Rey Vega, Sara Tressens |
IEEE Signal Process. Lett. | 1 |
| 2006 | Stereophonic acoustic echo cancellation: analysis of the misalignment in the frequency domainabstractThe performance in terms of misalignment of adaptive algorithms, in general, is dependent on the conditioning of the input signal covariance matrix. The performance of two-channel adaptive algorithms is further degraded by the high interchannel coherence between the two input signals. In this letter, we establish the relationship between interchannel coherence of the two input signals and condition of the corresponding covariance matrix for stereo acoustic echo cancellation application. We show how this relationship affects the misalignment of a frequency-domain adaptive algorithm. We provide simulation results for both white Gaussian noise and speech input to verify our mathematical analysis. Andy W. H. Khong, Jacob Benesty, Patrick A. Naylor |
IEEE Signal Process. Lett. | 2 |
| 2006 | Robust extended multidelay filter and double-talk detector for acoustic echo cancellationabstractWe propose an integrated acoustic echo cancellation solution based on a novel class of efficient and robust adaptive algorithms in the frequency domain, the extended multidelay filter (EMDF). The approach is tailored to very long adaptive filters and highly auto-correlated input signals as they arise in wideband full-duplex audio applications. The EMDF algorithm allows an attractive tradeoff between the well-known multidelay filter and the recursive least-squares algorithm. It exhibits fast convergence, superior tracking capabilities of the signal statistics, and very low delay. The low computational complexity of the conventional frequency-domain adaptive algorithms can be maintained thanks to efficient fast realizations. We also show how this approach can be combined efficiently with a suitable double-talk detector (DTD). We consider a corresponding extension of a recently proposed DTD based on a normalized cross-correlation vector whose performance was shown to be superior compared to other DTDs based on the cross-correlation coefficient. Since the resulting DTD also has an EMDF structure it is easy to implement, and the fast realization also carries over to the DTD scheme. Moreover, as the robustness issue during double talk is particularly crucial for fast-converging algorithms, we apply the concept of robust statistics into our extended frequency-domain approach. Due to the robust generalization of the cost function leading to a so-called M-estimator, the algorithms become inherently less sensitive to outliers, i.e., short bursts that may be caused by inevitable detection failures of a DTD. The proposed structure is also well suited for an efficient generalization to the multichannel case Herbert Buchner, Jacob Benesty, Tomas Gänsler, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | New insights into the noise reduction Wiener filterabstractThe problem of noise reduction has attracted a considerable amount of research attention over the past several decades. Among the numerous techniques that were developed, the optimal Wiener filter can be considered as one of the most fundamental noise reduction approaches, which has been delineated in different forms and adopted in various applications. Although it is not a secret that the Wiener filter may cause some detrimental effects to the speech signal (appreciable or even significant degradation in quality or intelligibility), few efforts have been reported to show the inherent relationship between noise reduction and speech distortion. By defining a speech-distortion index to measure the degree to which the speech signal is deformed and two noise-reduction factors to quantify the amount of noise being attenuated, this paper studies the quantitative performance behavior of the Wiener filter in the context of noise reduction. We show that in the single-channel case the a posteriori signal-to-noise ratio (SNR) (defined after the Wiener filter) is greater than or equal to the a priori SNR (defined before the Wiener filter), indicating that the Wiener filter is always able to achieve noise reduction. However, the amount of noise reduction is in general proportional to the amount of speech degradation. This may seem discouraging as we always expect an algorithm to have maximal noise reduction without much speech distortion. Fortunately, we show that speech distortion can be better managed in three different ways. If we have some a priori knowledge (such as the linear prediction coefficients) of the clean speech signal, this a priori knowledge can be exploited to achieve noise reduction while maintaining a low level of speech distortion. When no a priori knowledge is available, we can still achieve a better control of noise reduction and speech distortion by properly manipulating the Wiener filter, resulting in a suboptimal Wiener filter. In case that we have multiple microphone sensors, the multiple observations of the speech signal can be used to reduce noise with less or even no speech distortion. Jingdong Chen, Jacob Benesty, Yiteng Huang, Simon Doclo |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | A recursive estimation of the condition number in the RLS algorithm [adaptive signal processing applications]abstractThe recursive least-squares (RLS) algorithm is one of the most popular adaptive algorithms in the literature. This is due to the fact that it is easily derived and exactly solves the normal equations. In this paper, we present a very efficient way to recursively estimate the condition number of the input signal covariance matrix by utilizing fast versions of the RLS algorithm. We also quantify the misalignment of the RLS algorithm with respect to the condition number. Jacob Benesty, Tomas Gänsler |
ICASSP (4) | 1 |
| 2005 | Time delay estimation via multichannel cross-correlation [audio signal processing applications]abstractTime delay estimation (TDE) in a reverberant acoustical environment is a very challenging and difficult problem. This paper tackles the problem by exploiting the redundant information provided by multiple microphone sensors. To do so, the multichannel crosscorrelation coefficient (MCCC) is re-derived, in a new way, to connect it to the well-known linear interpolation technique. Some interesting properties and bounds of MCCC are discussed, and a recursive algorithm is then introduced so that MCCC can be estimated and updated efficiently when new data snapshots are available. We then apply the MCCC to the TDE problem, resulting a multichannel cross-correlation algorithm that can be treated as a natural generalization of the generalized cross-correlation (GCC) TDE method to the multichannel case. It is shown that this method can take advantage of the redundancy provided by multiple microphone sensors to improve TDE against both reverberation and noise. Jingdong Chen, Yiteng Huang, Jacob Benesty |
ICASSP (3) | 3 |
| 2005 | Adaptive blind SIMO identification: derivation of an optimal step size for the unconstrained multichannel LMS algorithmabstractAdaptive algorithms for blindly identifying SIMO systems are appealing because of their computational efficiency and capability of continuously tracking a time-varying system. Adaptive multichannel LMS (MCLMS) algorithms (with and without the unit-norm constraint) are analyzed and the optimal step size is derived. A simple yet effective variable step-size unconstrained MCLMS algorithm is proposed and its performance is evaluated with simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (3) | 2 |
| 2005 | Generalized multichannel frequency-domain adaptive filtering: efficient realization and application to hands-free speech communication
Herbert Buchner, Jacob Benesty, Walter Kellermann |
Signal Process. | 2 |
| 2005 | A generalized MVDR spectrumabstractThe minimum variance distortionless response (MVDR) approach is very popular in array processing. It is also employed in spectral estimation where the Fourier matrix is used in the optimization process. First, we give a general form of the MVDR where any unitary matrix can be used to estimate the spectrum. Second and most importantly, we show how the MVDR method can be used to estimate the magnitude squared coherence function, which is very useful in so many applications but so few methods exist to estimate it. Simulations show that our algorithm gives much more reliable results than the one based on the popular Welch's method. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Signal Process. Lett. | 1 |
| 2005 | Optimal step size of the adaptive multichannel LMS algorithm for blind SIMO identificationabstractAdaptive algorithms for blindly identifying single-input multiple-output (SIMO) systems are appealing because of their computational efficiency and capability of continuously tracking a time-varying system. Adaptive multichannel least-mean-square (MCLMS) algorithms (with and without the unit-norm constraint) are analyzed, and the optimal step size is derived. A simple yet effective variable step-size MCLMS algorithm is proposed, and its performance is evaluated with simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 2 |
| 2005 | A Blind Channel Identification-Based Two-Stage Approach to Separation and Dereverberation of Speech Signals in a Reverberant EnvironmentabstractBlind separation of independent speech sources from their convolutive mixtures in a reverberant acoustic environment is a difficult problem and the state-of-the-art blind source separation techniques are still unsatisfactory. The challenge lies in the coexistence of spatial interference from competing sources and temporal echoes due to room reverberation in the observed mixtures. Focusing only on optimizing the signal-to-interference ratio is inadequate for most if not all speech processing systems. In this paper, we deduce that spatial interference and temporal echoes can be separated and an M/spl times/N MIMO system will be converted into M SIMO systems that are free of spatial interference. Furthermore we show that the channel matrices of these SIMO systems are irreducible if the channels from the same source in the MIMO system do not share common zeros. Thereafter we can apply the Bezout theorem to remove reverberation in those SIMO systems. Such a two-stage procedure leads to a novel sequential source separation and speech dereverberation algorithm based on blind multichannel identification. Simulations with measurements obtained in the varechoic chamber at Bell Labs demonstrate the success and robustness of the proposed algorithm in highly reverberant acoustic environments. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | An exponentiated gradient adaptive algorithm for blind identification of sparse SIMO systemsabstractSparse impulse responses are encountered in many acoustic and wireless channels. Recently, a class of exponentiated gradient (EG) algorithms has been proposed. One of the algorithms belonging to this class, the so-called EG/spl plusmn/ algorithm, converges and tracks much better than the classical stochastic gradient, or LMS, algorithm for sparse impulse responses. We apply this technique to blind identification of a sparse SIMO system and develop the multichannel EG/spl plusmn/ algorithm. A simple experiment demonstrates its advantage in convergence compared to the MCLMS algorithm. Jacob Benesty, Yiteng Huang, Jingdong Chen |
ICASSP (2) | 1 |
| 2004 | An adaptive blind SIMO identification approach to joint multichannel time delay estimationabstractTime delay estimation (TDE) is a difficult problem in a reverberant environment and the traditional generalized cross-correlation (GCC) methods perform poorly. The adaptive eigenvalue decomposition (AED) algorithm recently proposed by the authors exploits a blind channel identification (BCI) technique and deals with room reverberation more effectively. The AED algorithm was developed for a two-channel system. It requires that the two channels do not share any common zeros (a necessary condition of system identifiability). We generalize the AED algorithm to multichannel (more than 2) systems. Compared to the AED algorithm, the generalized method is more robust since it is less likely for all channels to share a common zero when more sensors are used. Jingdong Chen, Yiteng Huang, Jacob Benesty |
ICASSP (4) | 3 |
| 2004 | Separating ISI and CCI in a two-step FIR Bezout equalizer for MIMO systems of frequency-selective channelsabstractThe use of multiple antennas at both the transmitter and the receiver in wireless communications implies a great channel capacity, but the signal detection is a crucial problem in achieving channel capacity, particularly in the general case of frequency selective channels, where both intersymbol interference (ISI) and cochannel interference (CCI) are significant. We show that ISI and CCI can be separated and then be cancelled in two different steps. We develop a two-step FIR Bezout equalizer and deduce the theoretically smallest length of its equalization filter. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (4) | 2 |
| 2004 | Time-delay estimation via linear interpolation and cross correlationabstractTime-delay estimation (TDE), which aims at measuring the relative time difference of arrival (TDOA) between different channels is a fundamental approach for identifying, localizing, and tracking radiating sources. Recently, there has been a growing interest in the use of TDE based locator for applications such as automatic camera steering in a room conferencing environment where microphone sensors receive not only the direct-path signal, but also attenuated and delayed replicas of the source signal due to reflections from boundaries and objects in the room. This multipath propagation effect introduces echoes and spectral distortions into the observation signal, termed as reverberation, which severely deteriorates a TDE algorithm in its performance. This paper deals with the TDE problem with emphasis on combating reverberation using multiple microphone sensors. The multichannel cross correlation coefficient (MCCC) is rederived here, in a new way, to connect it to the well-known linear interpolation technique. Some interesting properties and bounds of the MCCC are discussed and a recursive algorithm is introduced so that the MCCC can be estimated and updated efficiently when new data snapshots are available. We then apply the MCCC to the TDE problem. The resulting new algorithm can be treated as a natural generalization of the generalized cross correlation (GCC) TDE method to the multichannel case. It is shown that this new algorithm can take advantage of the redundancy provided by multiple microphone sensors to improve TDE against both reverberation and noise. Experiments confirm that the relative time-delay estimation accuracy increases with the number of sensors. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | A fast recursive algorithm for optimum sequential signal detection in a BLAST systemabstractBLAST (Bell Laboratories layered Space-Time) wireless systems are multiple-antenna communication schemes which can achieve very high spectral efficiencies in scattering environments, with no increase in bandwidth or transmitted power. The most popular and, by far, the most practical architecture is the so-called vertical BLAST (V-BLAST). The signal detection algorithm of a V-BLAST system is computationally very intensive. If the number of transmitters is M and is equal to the number of receivers, this complexity is proportional to M/sup 4/ at each sample time. In this paper, we propose a simple and very efficient algorithm that reduces the complexity by a factor of M. Jacob Benesty, Yiteng Huang, Jingdong Chen |
ICASSP (5) | 1 |
| 2003 | An extended multidelay filter: fast low-delay algorithms for very high-order adaptive systemsabstractWe propose a novel class of efficient adaptive algorithms in the frequency domain that is tailored to very long adaptive filters and highly autocorrelated input signals as they arise, e.g., in high-quality full-duplex audio applications. The approach exhibits good tracking capabilities of the signal statistics and very low delay. Moreover, it is shown that the low order of computational complexity of the conventional frequency-domain adaptive algorithms can be maintained thanks to efficient realizations. The algorithm allows a tradeoff between the well-known multidelay filter (MDF) and the recursive least-squares (RLS) algorithm. It is also well suited for an efficient generalization to the multichannel case. Herbert Buchner, Walter Kellermann, Jacob Benesty |
ICASSP (5) | 3 |
| 2003 | Robust time delay estimation exploiting spatial correlationabstractTo find the position of an acoustic source in a room, a set of relative delays among different microphone pairs has to be determined. The generalized cross-correlation method is the most popular to do so and is well explained in a landmark paper (Knapp and Carter (1996)). In this paper, we show how we can take advantage of the redundancy when more than two microphones are available. It is believed that the redundancy will help to better cope with noise and reverberation. The idea of cross-correlation coefficient between two signals is generalized to the multichannel case by using the notion of spatial prediction. The multichannel spatial correlation matrix is then deduced and it is shown how it can be used for time delay estimation. Jingdong Chen, Jacob Benesty, Yiteng Huang |
ICASSP (5) | 2 |
| 2003 | Adaptive blind identification of SIMO systems using channel cross-relation in the frequency domainabstractThe implementation of existing methods for blind identification of single-input multiple-output (SIMO) systems is limited in practice since they are difficult to execute in an adaptive mode and are, in general, computationally intensive. We extend our previous study (Huang, Y. and Benesty, J., Sig. Processing, vol.82, no.8, p.99-110, 2002) into the frequency domain and propose an unconstrained normalized multi-channel frequency-domain LMS (UNMCFLMS) algorithm. Numerical simulations show that the UNMCFLMS algorithm performs as well as (for a SIMO system with relatively short channel impulse responses) or better than (for a SIMO system with long channel impulse responses) its time-domain counterpart and the cross-relation (CR) batch method in practical situations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (6) | 2 |
| 2003 | Robust time delay estimation exploiting redundancy among multiple microphonesabstractTo find the position of an acoustic source in a room, typically, a set of relative delays among different microphone pairs needs to be determined. The generalized cross-correlation (GCC) method is the most popular to do so and is well explained in a landmark paper by Knapp and Carter. In this paper, the idea of cross-correlation coefficient between two random signals is generalized to the multichannel case by using the notion of spatial prediction. The multichannel spatial correlation matrix is then deduced and its properties are discussed. We then propose a new method based on the multichannel spatial correlation matrix for time delay estimation. It is shown that this new approach can take advantage of the redundancy when more than two microphones are available and this redundancy can help the estimator to better cope with noise and reverberation. Jingdong Chen, Jacob Benesty, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2002 | An improved PNLMS algorithmabstractRecently, the proportionate normalized least mean square (PNLMS) algorithm was developed for use in network echo cancelers. In comparison to the normalized least mean square (NLMS) algorithm, PNLMS has very fast initial convergence and tracking when the echo path is sparse. Unfortunately, when the impulse response is dispersive, the PNLMS converges much slower than NLMS. This implies that the rule proposed in PNLMS is far from optimal. In many simulations, it seems that we fully benefit from PNLMS only when the impulse response is close to a delta function. In this paper, we propose a new rule that is more reliable than the one used in PNLMS. Many simulations show that the new algorithm (improved PNLMS) performs better than NLMS and PNLMS, whatever the nature of the impulse response is. Jacob Benesty, Steven L. Gay |
ICASSP | 1 |
| 2002 | An adaptive nonlinearity solution to the uniqueness problem of stereophonic echo cancellationabstractThe major difference between two-channel and single-channel echo cancellation is the nonuniqueness problem in the two-channel case. In previous work, this nonuniqueness problem has been linked to the coherence between the two incoming audio channels. One proven solution to this problem is to distort the signals with a nonlinear device. In this work, we present an adaptive nonlinear device that incorporates new theoretical connections between the level of nonlinearity and the performance of the echo canceler. Adapting the level of nonlinearity is done in such a way that a pre-specified maximum misalignment is maintained while improving the perceived quality by minimizing the introduced distortion. Moreover, all the ideas presented can be generalized to the multichannel (> 2) case. Tomas Gänsler, Jacob Benesty |
ICASSP | 2 |
| 2002 | Adaptive blind channel identification: Multi-channel least mean square and Newton algorithmsabstractThe problem of identifying a single-input multiple-output FIR system without a training signal, the so-called blind system identification, is addressed and two adaptive multi-channel approaches, least mean square (LMS) and Newton algorithms, are proposed. In contrast to the existing batch blind channel identification schemes, the proposed algorithms construct an error signal based on the cross relations between different channels in a novel, systematic way. The corresponding cost (error) function is easy to manipulate and facilitates the use of adaptive filtering methods for an efficient blind channel identification scheme. It is theoretically shown and practically demonstrated by numerical studies that the proposed algorithms converge in the mean to the desired channel impulse responses for an identifiable system. Yiteng Huang, Jacob Benesty |
ICASSP | 2 |
| 2002 | Adaptive multi-channel least mean square and Newton algorithms for blind channel identification
Yiteng Huang, Jacob Benesty |
Signal Process. | 2 |
| 2002 | New insights into the stereophonic acoustic echo cancellation problem and an adaptive nonlinearity solutionabstractWe expand the knowledge regarding the problems of two-channel (or stereophonic) echo cancellation. The major difference between two-channel and the single-channel echo cancellation is the problem of nonunique solutions in the two-channel case. In previous work, this nonuniqueness problem has been linked to the coherence between the two incoming audio channels. One proven solution to this problem is to distort the signals with a nonlinear device. In this work, we present new theory that gives insight to the existing links between: (i) coherence and level of distortion, and (ii) coherence and achievable misalignment of the stereophonic echo canceler. Furthermore, we present an adaptive nonlinear device that incorporates this new knowledge in such a way that a pre-specified maximum misalignment is maintained while improving the perceived quality by minimizing the introduced distortion. Moreover, all the ideas presented can be generalized to the multichannel (>2) case. Tomas Gänsler, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | A robust fast recursive least squares adaptive algorithmabstractVery often, in the context of system identification, the error signal which is by definition the difference between the system and model filter outputs is assumed to be zero-mean, white, and Gaussian. In this case, the least squares estimator is equivalent to the maximum likelihood estimator and hence, it is asymptotically efficient. While this supposition is very convenient and extremely useful in practice, adaptive algorithms optimized on this may be very sensitive to minor deviations from the assumptions. We propose here to model this error with a robust distribution and deduce from it a robust fast recursive least squares adaptive algorithm (least squares is a misnomer here but convenient to use). We then show how to successfully apply this new algorithm to the problem of network echo cancellation combined with a double-talk detector. Jacob Benesty, Tomas Gänsler |
ICASSP | 1 |
| 2001 | Dynamic resource allocation for network echo cancellationabstractNetwork echo canceler chips are designed to handle several channels simultaneously. With the processing speeds now available, a single chip might handle several hundred channels. In current implementations, however, the adaptation algorithm is designed for a single channel, and the computations are replicated N/sub c/ times, where N/sub c/ is the number of channels. With such an implementation, the computational requirement is N/sub c/ times the peak load for a single channel. The number of computations required in each channel, however, varies widely over time. Therefore, a considerable reduction in computational load can be achieved by designing the system for the average load plus a margin to account for load variations. The reduction in complexity is achieved by exploiting three features: (a) the inherent pauses in conversations; (b) the sparseness of network echo paths; and (c) the fact that an adaptive filter does not need to be updated when the error signal is small. It is shown that, in principle, such a design can reduce the computational load by a very large factor - perhaps as large as thirty. It remains to be seen whether a customized hardware architecture can be implemented to take advantage fully of the proposed algorithm. Tomas Gänsler, Jacob Benesty, Man Mohan Sondhi, Steven L. Gay |
ICASSP | 2 |
| 2001 | A frequency-domain double-talk detector based on a normalized cross-correlation vector
Tomas Gänsler, Jacob Benesty |
Signal Process. | 2 |
| 2001 | A real-time implementation of a stereophonic acoustic echo cancelerabstractTeleconferencing systems employ acoustic echo cancelers to reduce echoes that result from the coupling between loudspeaker and microphone. To enhance the sound realism, two-channel audio is necessary. However, stereophonic acoustic echo cancellation (SAEC) is more difficult to solve because of the necessity to uniquely identify two acoustic paths, which becomes problematic since the two excitation signals are highly correlated. In this paper, a wideband stereophonic acoustic echo canceler is presented. The fundamental difficulty of stereophonic acoustic echo cancellation is described and an echo canceler based on a fast recursive least squares (FRLS) algorithm in a subband structure, with equidistant frequency bands, is proposed. The structure has been used in a real-time implementation, with which experiments have been performed. In this paper, simulation results of this implementation on real life recordings, with 8 kHz bandwidth, are studied. The results clearly verify that the theoretic fundamental problem of SAEC also applies in real-life situations. They also show that more sophisticated adaptive algorithms are needed in the lower frequency regions than in the higher regions. Peter Eneroth, Steven L. Gay, Tomas Gänsler, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 4 |
| 2001 | Real-time passive source localization: a practical linear-correction least-squares approachabstractA linear-correction least-squares estimation procedure is proposed for the source localization problem under an additive measurement error model. The method, which can be easily implemented in a real-time system with moderate computational complexity, yields an efficient source location estimator without assuming a priori knowledge of noise distribution. Alternative existing estimators, including likelihood-based, spherical intersection, spherical interpolation, and quadratic-correction least-squares estimators, are reviewed and comparisons of their complexity, estimation consistency and efficiency against the Cramer-Rao lower bound are made. Numerical studies demonstrate that the proposed estimator performs better under many practical situations. Yiteng Huang, Jacob Benesty, Gary W. Elko, Russell M. Mersereati |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Investigation of several types of nonlinearities for use in stereo acoustic echo cancellationabstractIn this paper, we investigate several types of nonlinearities used for the unique identification of receiving room impulse responses in stereo acoustic echo cancellation. The effectiveness is quantified by the mutual coherence of the transformed signals. The perceptual degradation is studied by psycho-acoustic experiments in terms of subjective quality and localization accuracy in the medial plane. The results indicate that, of the several nonlinearities considered, ideal half-wave rectification appears to be the best choice for speech. For music, the nonlinearity parameter of the ideal rectifier must be readjusted. The smoothed rectifier does not require this readjustment, but is a little more difficult to implement. Dennis R. Morgan, Joseph L. Hall, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 3 |
| 2000 | Frequency-domain adaptive filtering revisited, generalization to the multi-channel case, and application to acoustic echo cancellationabstractWe derive a new frequency-domain adaptive algorithm by using a frequency-domain recursive least squares criterion, minimizing an error signal in the frequency-domain. We then derive an exact adaptive algorithm from the so-called normal equation. It is shown that the obtained algorithm is complex to implement, and to reduce the complexity we need to remove a constraint resulting in the unconstrained frequency-domain LMS (UFLMS) algorithm. Most importantly, we generalize all this to the multi-channel case, thereby exploiting the cross-power spectra among all the channels which is very important (for a fast convergence rate) in multichannel acoustic echo cancellation (AEC), where the input signals are highly correlated. Jacob Benesty, Dennis R. Morgan |
ICASSP | 1 |
| 2000 | A robust proportionate affine projection algorithm for network echo cancellationabstractEcho cancelers which cover longer impulse responses (/spl ges/64 ms) are desirable. Long responses create a need for more rapidly converging algorithms in order to meet the specifications for network echo cancelers devised by the ITU (International Telecommunication Union). In general, faster convergence implies a higher sensitivity to near-end disturbances, especially "double-talk." Previously, a fast converging algorithm called the proportionate NLMS (normalized least mean squares) algorithm (PNLMS) has been proposed. This algorithm exploits the sparseness of the echo path in order to increase the convergence rate. A robust version of PNLMS has also been presented which combines a double-talk detector with techniques from robust statistics to make the algorithm insensitive to double-talk. This paper presents a generalization of the robust PNLMS algorithm to a robust proportionate aAffine projection algorithm (APA) called PAPA that converges very fast. Tomas Gänsler, Jacob Benesty, Steven L. Gay, Man Mohan Sondhi |
ICASSP | 2 |
| 2000 | Passive acoustic source localization for video camera steeringabstractA multi-input one-step least-squares (OSLS) algorithm for passive source localization is proposed. It is shown that the OSLS algorithm is mathematically equivalent to the so-called spherical interpolation (SI) method but with less computational complexity. The OSLS/SI method uses spherical equations (instead of hyperbolic equations) and solves them in a least-squares sense. Based on the adaptive eigenvalue decomposition time delay estimation method previously proposed by the same authors and the OSLS source localization algorithm, a real-time passive source localization system for video camera steering is presented. The system demonstrates many desirable features such as accuracy, portability, and robustness. Yiteng Huang, Jacob Benesty, Gary W. Elko |
ICASSP | 2 |
| 2000 | A new class of doubletalk detectors based on cross-correlationabstractA doubletalk detector (DTD) is used with an echo canceler to sense when far-end speech is corrupted by near-end speech. Its role is to freeze the adaptation of the model filter when near-end speech is present in order to avoid divergence of the adaptive algorithm. Several authors have proposed to use the cross-correlation coefficient vector between the input signal vector x and the scalar output y for a DTD. We show in this paper that this measure is not appropriate and propose a modified form that meets, in an optimal way, the needs for an efficient DTD. By extension, we also propose a definition of the normalized cross-correlation matrix between two vectors and show a link with the coherence function. Jacob Benesty, Dennis R. Morgan, Jun H. Cho |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Double-talk robust fast converging algorithms for network echo cancellationabstractThere is a need for echo cancelers for echo paths with long impulse responses (/spl ges/64 ms). This in turn creates a need for more rapidly converging algorithms in order to meet the specifications for network echo cancelers. Faster convergence, however, in general implies a higher sensitivity to near-end disturbances, especially "double-talk." Previously, a fast converging algorithm has been proposed called proportionate normalized least mean squares (PNLMS) algorithm. This algorithm exploits the sparseness of the echo path and has the advantage that no detection of active coefficients is needed. In this paper we propose a method for making the PNLMS algorithm more robust against double-talk. The slower divergence rate of these algorithms in combination with a standard Geigel double-talk detector improves the performance of a network echo canceler considerably during double-talk. The principle is based on a scaled nonlinearity which is applied to the residual error signal. This results in the robust PNLMS algorithm which diverges much slower than PNLMS and standard NLMS. Tradeoff between convergence and divergence rate is easily adjusted with one parameter and the added complexity is about seven instructions per sample which is less than 0.3% of the total load of a PNLMS algorithm with 512 filter coefficients. A generalization of the robust PNLMS algorithm to a robust proportionate affine projection algorithm (APA) is also presented. It converges very fast, and unlike PNLMS, is not as dependent on the assumption of a sparse echo path response. The complexity of the robust proportionate APA of order two is roughly the same as that of PNLMS. Tomas Gänsler, Steven L. Gay, Man Mohan Sondhi, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 4 |
| 1999 | A least squares component normalization approach to blind channel identificationabstractWe describe a new method for blind system identification that uses the cross relation properties between two or more sensor signals to estimate the impulse responses of the channels. The method performs as well or better than other similar blind identification techniques under noisy and ill-conditioned channel conditions, and is computationally simpler to implement. Carlos Avendaño, Jacob Benesty, Dennis R. Morgan |
ICASSP | 2 |
| 1999 | Synthesized stereo combined with acoustic echo cancellation for desktop conferencingabstractOne promising application in communications is desktop conferencing, which can involve several participants over a widely distributed area. Synthesized stereophonic sound will enable a listener to spatially separate one remote talker from another and thereby improve understanding. In such a scenario, we assume we are located in a hands-free environment where the composite acoustic signal is presented over loudspeakers, thus requiring acoustic echo cancellation. In this paper, we explain some of the methods that can be used to synthesize stereo sound and how such methods can be combined efficiently with stereo acoustic echo cancellation in the face of several difficult problems. Jacob Benesty, Dennis R. Morgan, Joseph L. Hall, Man Mohan Sondhi |
ICASSP | 1 |
| 1999 | Adaptive eigenvalue decomposition algorithm for real time acoustic source localization systemabstractTo locate an acoustic source in a room, the relative delay between microphone pairs must be determined efficiently and accurately. However, most traditional time delay estimation (TDE) algorithms fail in reverberant environments. A new approach is proposed that takes into account the reverberation of the room. A real time PC-based TDE system running under Microsoft/sup TM/ Windows system was developed with three TDE techniques: classical cross-correlation, phase transform, and a new algorithm that is proposed in this paper. The system provides an interactive platform that allows users to compare performance of these algorithms. Yiteng Huang, Jacob Benesty, Gary W. Elko |
ICASSP | 2 |
| 1999 | An objective technique for evaluating doubletalk detectors in acoustic echo cancelersabstractEcho cancelers commonly employ a doubletalk detector (DTD), which is essential to keep the adaptive filter from diverging in the presence of near-end speech and other disruptive noise. There have been numerous algorithms to detect doubletalk in an acoustic echo canceler (AEC). In those applications, typically, the threshold is chosen only by some heuristic method and the performance evaluation is very subjective. In this study, we develop a way to objectively evaluate DTD algorithms based on the standard statistical methods of detection theory. A receiver operating characteristic (ROC) is derived to characterize DTD performance. Several DTD algorithms are examined and simulated under typical real-world operating conditions using measured room responses and signals taken from a digital speech database. The DTD methods are then evaluated and compared using the ROC metric. Jun H. Cho, Dennis R. Morgan, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 3 |
| 1998 | Stereophonic acoustic echo cancellation using nonlinear transformations and comb filteringabstractStereophonic sound becomes more and more important in a growing number of applications (such as teleconferencing, multimedia workstations, televideo gaming, etc.) where spatial realism is demanded. Such hands-free systems need stereophonic acoustic echo cancelers (AECs) to reduce echos that result from coupling between loudspeakers and microphones in full-duplex communication. We propose a new stereo AEC based on two experimental observations: (a) the stereo effect is due mostly to sound energy below about 1 kHz and (b) comb filtering above 1 kHz does not degrade auditory localization. The principle of the proposed structure is to use one stereo AEC at low frequencies (e.g. below 1 kHz) with nonlinear transformations on the input signals and another stereo AEC at higher frequencies (e.g. above 1 kHz) with complementary comb filters on the input signals. Jacob Benesty, Dennis R. Morgan, Joseph L. Hall, Man Mohan Sondhi |
ICASSP | 1 |
| 1998 | On the evaluation of estimated impulse responsesabstractWe discuss the philosophy of evaluating estimated impulse responses for applications in which the overall scaling or gain is irrelevant, as, for example, in blind channel identification. We argue that no matter how the estimate is obtained, the performance should be independently evaluated using an error measure that is appropriate for a problem of this class. Of several possible error measures, the normalized projection error, which is obtained by minimizing the l/sub 2/ norm over all possible gain values, seems to be the most natural and consistent. An example demonstrates that using an inappropriate error measure can produce misleading results. Dennis R. Morgan, Jacob Benesty, Man Mohan Sondhi |
IEEE Signal Process. Lett. | 2 |
| 1998 | A better understanding and an improved solution to the specific problems of stereophonic acoustic echo cancellationabstractTeleconferencing systems employ acoustic echo cancelers to reduce echoes that result from coupling between the loudspeaker and microphone. To enhance the sound realism, two-channel audio is necessary. However, in this case (stereophonic sound) the acoustic echo cancellation problem is more difficult to solve because of the necessity to uniquely identify two acoustic paths. We explain these problems in detail and give an interesting solution which is much better than previously known solutions. The basic idea is to introduce a small nonlinearity into each channel that has the effect of reducing the interchannel coherence while not being noticeable for speech due to self masking. Jacob Benesty, Dennis R. Morgan, Man Mohan Sondhi |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | A hybrid mono/stereo acoustic echo cancelerabstractIn many applications, such as teleconferencing, multimedia workstations, televideo gaming, etc., stereo sound is already, or will soon be, implemented to give spatial realism that mono systems cannot offer. In such hands-free systems, stereophonic acoustic echo cancelers are absolutely necessary for full-duplex communication. We propose a new acoustic echo canceler (AEC) based on a fundamental experimental observation that the stereo effect is due mostly to sound energy below about 1000 Hz. The principle of the hybrid mono/stereo AEC is to use stereophonic sound with a stereo AEC at low frequencies (e.g., below 1000 Hz) and monophonic sound with a conventional mono AEC at higher frequencies (e.g., above 1000 Hz). This solution is a good compromise between the complexity of a full-band stereo AEC and spatial realism. For the stereo case, we borrow from a previous innovation and add a small nonlinearity into each channel in order to accurately identify the two receiving room impulse responses. Jacob Benesty, Dennis R. Morgan, Man Mohan Sondhi |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | A better understanding and an improved solution to the problems of stereophonic acoustic echo cancellationabstractTeleconferencing systems employ acoustic echo cancelers (AECs) to reduce echos that result from coupling between the loudspeaker and microphone. To enhance the sound realism, two-channel audio is necessary. However, in this case (stereophonic sound) the acoustic echo cancellation problem is more difficult to solve because of the necessity to uniquely identify two acoustic paths. We explain these problems in detail and give an interesting solution which is much better than previously known solutions. The basic idea is to introduce a small nonlinearity into each channel that has the effect of reducing the interchannel coherence while not being noticeable for speech due to self masking. Jacob Benesty, Dennis R. Morgan, Man Mohan Sondhi |
ICASSP | 1 |
| 1996 | A fast two-channel projection algorithm for stereophonic acoustic echo cancellationabstractWe propose a new fast projection algorithm for stereophonic acoustic echo cancellation. This algorithm can be viewed as a generalization of the extended two-channel LMS algorithm which takes into account the correlation between the input signals. Moreover, this algorithm fits naturally with the framework of affine projection techniques extended to the two-channel case. Its computational complexity is less than half the complexity of the fastest two-channel RLS versions. Simulation results show the obtained performance. Fabrice Amand, Jacob Benesty, André Gilloire, Yves Grenier |
ICASSP | 2 |
| 1996 | A multichannel affine projection algorithm with applications to multichannel acoustic echo cancellationabstractA straightforward generalization of the so-called affine projection algorithm (APA) to the multichannel (MC) case is easily obtained. However, due to the strong correlation between the input signals of the various channels, the resulting algorithm converges very slowly. The article describes the way to overcome this problem and derives an efficient algorithm that makes use of additional orthogonal projections. Jacob Benesty, Pierre Duhamel, Yves Grenier |
IEEE Signal Process. Lett. | 1 |
| 1995 | Adaptive filtering algorithms for stereophonic acoustic echo cancellationabstractIt is likely that stereophonic (and more generally, multichannel) sound pick-up, transmission and diffusion will be implemented in future teleconference systems to provide the users with enhanced quality. Therefore, adequate solutions must be found to solve the problem of stereophonic acoustic echo which will occur in such systems. We explain in this paper the difference between the mono and two-channel systems and the behavior of the two-channel classical adaptive algorithms in comparison with the same algorithms in the mono-channel case. Also, we outline a new NLMS-like algorithm derived from the two-channel RLS algorithm as a first member of a family of improved two-channel adaptive filters. Jacob Benesty, Fabrice Amand, André Gilloire, Yves Grenier |
ICASSP | 1 |
| 1992 | A gradient-based adaptive algorithm with reduced complexity, fast convergence and good tracking characteristicsabstractA new block processing algorithm which requires a smaller number of arithmetic operations than the least mean square (LMS) algorithm, while converging faster is described. Theoretical conditions for this convergence to be faster are derived and checked by simulation. Some simulations are provided showing that, at least under some specific conditions, its tracking characteristics are also improved. These performances are obtained by grouping the computations corresponding to a block of successive inputs. Good performances are obtained even with very small blocks.> Jacob Benesty, Li Sheng Wen, Pierre Duhamel |
ICASSP | 1 |
| 1990 | A fast exact least mean square adaptive algorithmabstractA general block-formulation is presented for the LMS (least-mean-square) algorithm for adaptive filtering. This formulation has an exact equivalence with the initial LMS, hence retaining the same convergence properties while allowing a reduction in the arithmetic complexity, even for very small block lengths. Furthermore, tradeoffs between number of operations and convergence rate are obtainable by applying certain approximations to a matrix involved in the algorithm. The usual block LMS (BLMS) hence appears as one of the possible approximations, which explains some of its properties.> Jacob Benesty, Pierre Duhamel |
ICASSP | 1 |