VLDB 2026 Research / reviewers in the wild / expert
Gongping Huang
dblp:148/9904
· DBLP profile ↗
61ranked-venue papers
18as first author
42since 2021 · last 2026
0000-0002-6825-7473ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 9 first-author · 29 since 2021Artificial intelligence and machine learning · 22 · 9 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference OptimizationabstractCurrent video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical limitations. First, existing methods face challenges in precisely aligning intricate semantic information together with rapid dynamic features. Second, foundational training lacks quantitative preference optimization for semantic-temporal alignment and audio quality. As a result, it fails to enhance integrated generation quality in cluttered multi-event scenes. To address these core limitations, this study proposes a novel V2A framework: MultiSoundGen. It introduces direct preference optimization (DPO) into the V2A domain, leveraging audio-visual pretraining (AVP) to enhance performance in complex multi-event scenarios. Our contributions include two key innovations: the first is SlowFast Contrastive AVP (SF-CAVP), a pioneering AVP model with a unified dual-stream architecture. SF-CAVP explicitly aligns core semantic representations and rapid dynamic features of audio-visual data to handle multi-event complexity; second, we integrate the DPO method into V2A task and propose AVP-Ranked Preference Optimization (AVP-RPO). It uses SF-CAVP as a reward model to quantify and prioritize critical semantic-temporal matches while enhancing audio quality. Experiments demonstrate that Multi-SoundGen achieves state-of-the-art (SOTA) performance in multi-event scenarios, delivering comprehensive gains across distribution matching, audio quality, semantic alignment, and temporal synchronization. Jianxuan Yang, Lipan Zhang, Xinyue Guo 0001, Gongping Huang |
ICMR | 6 |
| 2026 | Design of Low-Rank differential beamformers with constrained directivity or robustness
Kunlong Zhao, Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
Signal Process. | 4 |
| 2026 | Correlation maximization-based sampling rate offset estimation with multichannel node information
Yingke Zhao, Xueqin Luo, Jilu Jin, Gongping Huang, Haiyang Yao |
Signal Process. | 4 |
| 2026 | A third-order tensor decomposition based algorithm for speech dereverberation
Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
Signal Process. | 2 |
| 2026 | Revisiting Steering Limitations of LCMV Beamforming for Circular Microphone Arrays and a Mainlobe-Controlled SolutionabstractThis paper investigates the performance limitations of conventional linearly constrained minimum variance (LCMV) beamformers implemented with circular microphone arrays. In particular, we show that imposing null constraints on interference directions can lead to a deviation of the mainlobe from the desired steering angle. A theoretical analysis is presented to characterize this deviation, and a closed-form expression for the deviation is derived to reveal the underlying low-frequency behavior of LCMV beamformers. To address this issue, we propose a mainlobe-controlled LCMV (MC-LCMV) beamformer. Simulation results demonstrate that the proposed method substantially improves spatial directivity and speech enhancement performance compared with conventional LCMV beamformers. Wei Liu 0177, Gongping Huang, Jilu Jin, Xueqin Luo, Shoji Makino |
IEEE Signal Process. Lett. | 2 |
| 2026 | MCFLOW-SE: Efficient One-Step Multichannel Speech Enhancement via MeanflowabstractGenerative models have recently shown strong potential for high-fidelity speech enhancement but suffer from high inference latency due to iterative sampling. In contrast, discriminative models offer fast processing but tend to introduce non-linear distortions that degrade downstream task performance. To bridge this gap, we propose MCFLOW-SE, a spacing-aware generative framework enabling one-step multichannel speech enhancement. By learning an average velocity field that directly maps noisy multichannel observations to clean speech, MCFLOW-SE enables fast one-step inference while effectively leveraging spatial information across channels. Furthermore, sensor spacing is incorporated as positional encoding into the generative process, improving robustness across the evaluated array-spacing configurations. Experiments on both simulated and real-world datasets demonstrate that MCFLOW-SE delivers substantially faster inference than prior generative approaches, while achieving superior Automatic Speech Recognition (ASR) accuracy over discriminative baselines. These results support its potential to bridge the gap between high-fidelity generation and real-time applicability. Gongping Huang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Advances in Microphone Array Processing and Multichannel Speech EnhancementabstractThis paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas. The paper examines foundational developments in microphone array design and optimization, showcasing innovations that improved sound acquisition and enhanced speech intelligibility in noisy and reverberant environments. It then introduces recent advancements and cutting-edge research in the field, particularly the integration of deep learning techniques such as all-neural beamformers. The paper also explores critical applications, discussing their evolution and current state-of-the-art technologies that significantly impact user experience. Finally, the paper outlines future research directions, identifying challenges and potential solutions that could drive further innovation in these fields. By providing a comprehensive overview and forward-looking perspective, this paper aims to inspire ongoing research and contribute to the sustained growth and development of microphone arrays and multichannel speech enhancement. Gongping Huang, Jesper Rindom Jensen, Jingdong Chen, Jacob Benesty, Mads Græsbøll Christensen, Akihiko Sugiyama, Gary W. Elko, Tomas Gänsler |
ICASSP | 1 |
| 2025 | Microphone Array Beamforming for Speech Enhancement Based on Dynamic Mode DecompositionabstractMicrophone array beamforming is widely used to extract desired speech signals from noisy environments. While most research in this area focuses on utilizing spatial information, less attention is given to the intrinsic physical mechanisms underlying microphone array observations. This paper aims to address this gap by exploring these underlying factors through dynamic mode decomposition (DMD). Our contributions are twofold. 1) We develop a DMD-based signal model for microphone arrays to capture the relationships between observation signals at adjacent microphones. 2) We introduce a DMD-based preprocessing method and a corresponding beamforming approach based on this model. Simulation results show that our proposed method significantly enhances performance compared to conventional beamforming techniques. Wei Liu 0177, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2025 | On the Design of Low-Rank Differential Beamformers with Nonuniform Linear Microphone ArraysabstractKronecker product beamforming is an effective technique for designing beamformers with nonuniform linear arrays (NULAs). However, current techniques are restricted to NULAs with specific configurations, where the steering vector of the array is represented as a Kronecker product of steering vectors from smaller virtual arrays. This paper overcomes these constraints by proposing a novel approach to designing Kronecker product beamformers for NULAs from a low-rank perspective. Our approach involves decomposing the NULA into overlapping subarrays and organizing the sensor signals from these subarrays into a matrix. We then apply filters to both sides of this matrix to produce an output, which is then converted into a low-rank beamforming process. This method is highly adaptable and can be utilized for NULAs with any number of microphones. Hanchen Pei, Gongping Huang, Jilu Jin, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2025 | Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech EnhancementabstractSuperdirective beamformers are highly effective at suppressing directional interference and diffuse noise, but their practical use is often constrained by the problem of white noise amplification. Robust superdirective beamforming methods typically address this by imposing a constraint on the white noise gain (WNG). However, determining the appropriate WNG threshold in varying noise environments remains unclear. This paper introduces a data-driven approach to estimating the optimal WNG threshold. Subsequently, a more versatile and robust superdirective beamformer is developed by solving a quadratic eigenvalue problem (QEP). Experimental results show that this method outperforms traditional superdirective beamformers, which rely on a WNG threshold set through a fixed search range. Importantly, this approach functions as a distortionless beamformer, maintaining high fidelity of the desired acoustic signal and allowing for additional post-filtering if required. Hanchen Pei, Gongping Huang, Jilu Jin, Zhizheng Wu 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2025 | Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech EnhancementabstractSuperdirective beamformers, used with small microphone arrays, are highly attractive due to their high directivity and frequency-invariant beampatterns, making them well-suited for processing broadband acoustic and speech signals. However, these beamformers are very sensitive to array imperfections such as sensor mismatches and self-noise. To improve robustness, robust superdirective (RSD) beamformers have been developed, employing techniques such as diagonal loading or white-noise-gain constraints during their derivation. Although RSD beamformers offer enhanced robustness compared to classical superdirective beamformers, they cannot achieve the maximum directivity factor and lose some frequency-invariant properties, resulting in a beamwidth that is wider at low frequencies and narrower at high frequencies. As a result, RSD beamformers do not fully meet the criteria of true superdirective beamformers, providing less effective noise reduction and introducing some speech distortion. Post-filtering methods have been developed to improve noise reduction after RSD beamforming, but they often fail to address the distortion issues, especially when the speech source deviates from the array’s look direction. To overcome this limitation, this paper proposes a joint optimization approach that combines post-filtering with RSD beamformers. By using the output of RSD beamformers as input data and considering various deviations in look directions and array mismatches, we train a post-filtering network to further enhance the beamformer’s output. Experimental results on speech enhancement demonstrate the effectiveness and robustness of the proposed method. Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2025 | DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone ArraysabstractDirection-of-arrival (DOA) estimation is a key process in microphone array systems. The steered response power-based minimum variance distortionless response (SRP-MVDR) method performs very well in challenging acoustic environments but suffers from exponential complexity as the number of microphones increases. To improve the efficiency of SRP-MVDR for real-time applications, we propose a Kronecker product-based SRP-MVDR (SRP-KPMVDR) method designed for large rectangular microphone arrays. This approach begins with a rank-one approximation that represents the signal covariance matrix of a rectangular microphone array in Kronecker product form, which is essential for SRP-MVDR estimation. By utilizing the Kronecker product properties, the complex matrix inversion in SRP-MVDR is simplified to the inversion of two smaller matrices, significantly reducing computational complexity. Simulation results show that the SRP-KPMVDR method achieves comparable performance to the traditional SRP-MVDR while greatly decreasing the computational demands. Yichen Zeng, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2025 | LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band AttentionabstractDeep learning based end-to-end multi-channel speech enhancement methods have achieved impressive performance by leveraging sub-band, cross-band, and spatial information. However, these methods often demand substantial computational resources, limiting their practicality on terminal devices. This paper presents a lightweight multichannel speech enhancement network with decoupled fully connected attention (LMFCA-Net). The proposed LMFCA-Net introduces time-axis decoupled fully-connected attention (T-FCA) and frequency-axis decoupled fully-connected attention (F-FCA) mechanisms to effectively capture long-range narrow-band and cross-band information without recurrent units. Experimental results show that LMFCA-Net performs comparably to state-of-the-art methods while significantly reducing computational complexity and latency, making it a promising solution for practical applications. Yaokai Zhang, Hanchen Pei, Wanqi Wang, Gongping Huang |
ICASSP | 4 |
| 2025 | Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar GeometryabstractDifferential microphone arrays (DMAs) have garnered significant attention in recent research and development due to their high directivity and frequency-invariant beampatterns. However, DMAs frequently encounter substantial white noise amplification, which limits their practical applications. This paper addresses this issue by introducing a general method for designing robust DMAs with microphone arrays of arbitrary planar topology. The proposed approach approximates the beampattern using the Jacobi-Anger series expansion and constrains the white noise gain (WNG) to a specified value. This minimizes the error between the beampattern and the ideal directivity pattern while ensuring a reasonable level of robustness. A closed-form solution for the robust differential beamformer filter is derived using the quadratic eigenvalue problem (QEP) method. Simulation results demonstrate the feasibility and effectiveness of the proposed approach. Kunlong Zhao, Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2025 | TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
Ningning Pan, Gongping Huang |
INTERSPEECH | 4 |
| 2025 | On the Design of a Robust Superdirective Beamformer and Topology Parameter Optimization with Frustum-Shaped Microphone Arrays Featuring Multiple Rings
Kunlong Zhao, Gongping Huang, Jingdong Chen, Jacob Benesty, Zoran Cvetkovic |
INTERSPEECH | 2 |
| 2025 | Robust Fusion of Differential Beamformers for Speech Enhancement in Dynamic Interference ConditionsabstractDifferential microphone arrays are widely used for far-field sound acquisition due to their high directivity and compact geometry. However, they lack the flexibility to adapt in dynamic acoustic environments with multiple or moving interferers. This paper proposes a novel method for fusing multiple differential beamformers to improve robustness under such conditions. A set of beamformers is designed with distortionless constraints in the target direction and nulls in various potential interference directions. An online fusion strategy is then applied, where a subset of beamformer outputs is selected and adaptively combined at each time frame based on the criterion of minimizing the instantaneous output variance. Simulation results demonstrate that the proposed method achieves superior interference suppression and speech quality, while maintaining low computational complexity suitable for real-time processing. Kunlong Zhao, Xueqin Luo, Jilu Jin, Danqi Jin, Gongping Huang |
IEEE Signal Process. Lett. | 5 |
| 2024 | Beamforming Through Online Convex Combination of Differential BeamformersabstractThanks to their high directivity, compact size, and reliable performance, differential microphone arrays (DMAs) have attracted great interest from both industry and academia as they have demonstrated great potential to be used in a wide range of applications for high-fidelity speech acquisition. Nevertheless, in many real-world applications, DMAs powered with fixed differential beamformers are often inadequate in suppressing interference, particularly in environments with multiple or moving sources. To address this issue, this work develops an adaptive convex combination (ACC)-based method, which combines multiple differential beamformers in an online manner for enhanced performance. While the major contribution is a new real-time processing algorithm that facilitates optimal linear combinations of different differential beamformers, making them adapted to dynamic environments, the presented method also provides valuable insights as how to combine different beamformers for online robust implementation. Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2024 | On the Design of Planar Differential Microphone Arrays with Specified Beamwidth or Sidelobe LevelabstractThis paper investigates the problem of designing differential beam-formers with planar microphone arrays to achieve not only the desired target directivity pattern but also control the beamwidth (BW) or sidelobe level (SLL). We first discuss the target directivity patterns and express the Dolph-Chebyshev polynomial based form of target directivity patterns into linear combination of cylindrical harmonics. We then address the problem of designing differential beamformers through beampattern approximation based on the Jacobi-Anger series expansion. Two methods are subsequently developed: the first one involves designing beamformers to achieve the target directivity pattern while minimizing SLL under a pre-specified value of BW and the second one aims to attain the target directivity pattern while minimizing the null-to-null BW under a pre-specified level of SLL. Simulations are carried out to validate the method and the results demonstrate the properties of the proposed method. Xueqin Luo, Jilu Jin, Gongping Huang, Yingke Zhao, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2024 | Differential Beamforming with Null Constraints for Spherical Microphone ArraysabstractDifferential microphone arrays (DMAs) can measure both the acoustic pressure field and the differential acoustic pressure fields, which gives them great advantages in a wide range of applications for acoustic and speech signal acquisition. The core component of DMAs is the so-called differential beamformer, the design of which typically involves taking into account the a priori knowledge about the array geometry and the desired directivity pattern that is related to the differential sound field to respond. This paper deals with the design of differential beamformers with spherical microphone arrays. It presents a novel design approach based on the null constraints formed from the desired directivity pattern. In comparison with the exiting methods, the proposed approach only requires the information of the zeros in the beampattern, which provides notable flexibility and convenience for spherical DMA design in practical applications. Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2024 | Design of Fully Steerable Differential Beamformers With Linear SuperarraysabstractLinear differential microphone arrays (LDMAs) are commonly integrated into thin and portable devices to achieve high-fidelity speech acquisition. Traditional LDMAs typically consist of only omnidirectional microphones, which impose limitations on their ability to produce steerable spatial responses due to constraints in array element directivity and linear array geometry. A recent solution to this limitation involves integrating both omnidirectional and bidirectional microphones in LDMA design, enabling the creation of steerable spatial responses. This paper extends the core idea of integrating omnidirectional and bidirectional microphones, and develops a more general and comprehensive theory and method for designing steerable LDMAs. It makes two main contributions. Firstly, it introduces a general approach to designing steerable LDMAs, in which any type of directional microphones can be used. Secondly, it gives the minimum number of omnidirectional and directional microphones required to achieve a specific order of steerable LDMA. Simulations validate the proposed method and illustrate how omnidirectional and directional sensors can be combined to form the desired LDMAs. Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech DereverberationabstractDereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those methods generally assume that there is only a single speaker in the acoustic environment and, consequently, they suffer from significant performance degradation if multiple speakers participate in the conversation. How to deal with reverberation in multiple-speaker scenarios is still a challenging problem, which is studied in this work. We present a switching multichannel linear prediction filtering method, which designs multiple linear filters with each tracking one speaker. When some speaker is active, the corresponding filter and the weighted cross-correlation matrix are updated while the other filters are kept unchanged. To further improve the performance and reduce complexity, we apply the Kronecker product to decompose every linear prediction filter into a Kronecker product of two shorter filters: one is time-invariant and the other is time-varying. The former is estimated with a batch method (using only a few seconds of speech signal when the corresponding speaker starts to talk in the entire conversation) while a recursive least-squares algorithm is derived for identifying the time-varying set of Kronecker filters. Gongping Huang, Jacob Benesty, Israel Cohen, Emil Winebrand, Jingdong Chen, Walter Kellermann |
ICASSP | 1 |
| 2023 | Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function ModelabstractSpatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the performance of those methods degrades quickly as the reverberation increases. The underlying reason is that those methods are derived based on the multiplicative transfer function model with a rank-1 assumption, which does not hold true if reverberation is strong. To circumvent this issue, this paper proposes to use the convolutive transfer function (CTF) model to improve the source extraction performance and develop a spatially informed IVA algorithm. Simulations demonstrate the efficacy of the developed method even in highly reverberant environments. Xianrui Wang, Andreas Brendel, Gongping Huang, Yichen Yang 0010, Walter Kellermann, Jingdong Chen |
ICASSP | 3 |
| 2023 | Dimensionality Reduction of Room Acoustic Impulse Responses and Applications to System IdentificationabstractA room Acoustic Impulse Response (RAIR), which represents the sound propagation channel via direct and reflection paths from a source position to a microphone, plays a leading role in a broad range of acoustic signal processing applications, e.g., echo cancellation. In practical acoustic environments, it is not uncommon that an RAIR may consist of hundreds or even thousands of coefficients, making it challenging to identify and handle. This paper investigates the RAIR dimensionality reduction problem inspired from the concepts of dynamic mode decomposition. The objective is to find effective lower-dimensional representations of RAIRs, which are easier and more robust to identify and equalize. There are two main contributions of this work. First, we present an RAIR dimensionality reduction method. Second, we show how to apply this technique to the problem of acoustic system identification. Simulation results demonstrate that the proposed method is able to improve significantly the performance of acoustic system identification. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 1 |
| 2023 | Differential Beamforming From a Geometric PerspectiveabstractDifferential microphone arrays (DMAs) have demonstrated a great potential for solving the high-fidelity sound acquisition problem in a wide range of applications as they possess many good properties such as frequency-independent beampatterns with high directivity. A significant number of efforts have been devoted to the design of DMAs and the associated beamformers. As a result, many different types of DMAs and differential beamforming methods have been developed over the last few decades, some of which have been successfully deployed in real systems and commercial products. However, given an application, how to design a DMA to achieve optimal performances is still an open issue. This work studies the problem of designing linear DMAs (LDMAs) from a geometric perspective. Based on the fundamental observation that most practical and interesting DMA beampatterns have nulls in some directions, we define a criterion based on the orthogonality between the beamforming filter and the steering vector in the nulls' directions. We then derive a family of differential beamformers by optimizing the defined criterion, some of which are well known but derived from a different perspective, while others are new. Simulations and experiments are carried out, and the results validate the proposed method and developed differential beamformers. Jilu Jin, Jacob Benesty, Jingdong Chen, Gongping Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Design of Maximum Directivity Beamformers With Linear Acoustic Vector Sensor ArraysabstractThis paper studies the design of maximum directivity factor (MDF) beamformers based on uniform linear arrays (ULAs) consisting of acoustic vector sensors (AVSs). We first derive the main lobe constraints, which ensure that the beamformer's beampattern achieves a maximum in the look direction, and prove that any beamformer that satisfies the proposed constraints can be written as the sum of two orthogonal beamformers: the maximum white noise gain (MWNG) beamformer and a reduced-rank beamformer. Then, we derive the MDF beamformer by maximizing the directivity factor (DF) under the deduced constraints. We also derive a robust version of the MDF beamformer, which can keep the WNG above a pre-specified level. Compared to the conventional MDF beamformer based on ULAs with omnidirectional microphones, the designed MDF beamformer with uniform linear AVS arrays (ULAVSAs) can steer the beampattern to any look direction in the 3-dimensional space and achieves a higher directivity. The proposed MDF beamformer also outperforms the two-step MDF beamformer with ULAVSAs since it maximizes the DF. The proposed methods are validated through simulations as well as real experiments. Xueqin Luo, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty, Wen Zhang 0002, Mengyao Zhu 0003, Chunjian Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Design of 2D and 3D Differential Microphone Arrays With a Multistage FrameworkabstractDifferential microphone arrays (DMAs) have demonstrated a great potential for high-fidelity acoustic and speech signal acquisition in a wide range of applications since such arrays are able to achieve frequency-invariant beampatterns with high directivity. Consequently, a great number of efforts have been devoted to the design of DMAs and the associated beamformers in the literature. However, most of the methods only work for arrays with particular topologies, e.g., linear, circular, concentric circular, and spherical ones. How to design general two-dimensional (2D) and three-dimensional (3D) DMAs that can measure the desired differential sound field and form the desired spatial response in the 3D space remains an unsolved problem. This paper investigates this problem and presents a multistage design approach. The major contributions of this work are as follows. First, we reexamine the differentials of the acoustic pressure field in the 3D space and derive the general expression of the directivity patterns resulting from the spatial differential operation, which serves as the foundation for differential beamforming with 2D or 3D microphone arrays. Second, we present a multistage approach to the design of 2D and 3D DMAs, and deduce the relationship between the global beamformer and the beamformers at different stages as well as the relationship between their beampatterns. Third, several algorithms are presented for the design of differential as well as robust differential beamformers in this multistage framework. Simulation results validate the proposed approach and justifies its properties. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Fundamental Approaches to Robust Differential Beamforming With High Directivity FactorsabstractDifferential beamforming, which measures the spatial derivatives of the acoustic pressure field, can be used in a wide range of small devices that require high-fidelity sound and speech acquisition as it can achieve frequency-invariant spatial responses with high directivity factors (DFs). Since a differential process is inherently sensitive to sensors' self noise and other array imperfections, the most challenging problem in the design of any differential beamformer is how to achieve the maximum possible DF while maintaining a proper level of robustness for practical usage. While significant efforts have been made on this topic, the problem remains unsolved and further study is indispensable. This paper is devoted to dealing with this challenging problem. It presents a study on theory and methods to achieve the optimal and fundamental compromise between the white noise gain (WNG), which quantifies how robust is the beamformer, and the DF in differential beamforming. The major contributions of this work are as follows. 1) We show and prove that any null constrained fixed beamformer can be decomposed as the sum of two orthogonal filters, i.e., the maximum WNG (MWNG) beamformer and a reduced-rank one. Based on this decomposition, we develop three kinds of differential beamformers from the WNG perspective, which can achieve a flexible and optimal compromise between DF and WNG. 2) We show that a transformed null constrained beamformer can also be decomposed as the sum of two orthogonal filters, i.e., the transformed maximum DF (MDF) beamformer and another reduced-rank one. Based on this decomposition, we also develop three kinds of differential beamformers, which can obtain the desired level of DF while using the rest of the degrees of freedom to maximize the WNG. Simulations are performed to validate the theoretical analysis and developed differential beamformers. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Kronecker Product Multichannel Linear Filtering for Adaptive Weighted Prediction Error-Based Speech DereverberationabstractReverberation, whichis caused by late reflections, impairs not only speech quality but also intelligibility. Consequently, dereverberation, a process to mitigate the impact of reverberation, has attracted significant research interests. Numerous approaches have been developed in the literature, among which the weighted-prediction-error (WPE) one has demonstrated promising potential for reducing or eliminating reverberation. The WPE method has been well studied and several variants have been developed. The adaptive one, called adaptive WPE (AWPE) method, has been widely investigated for use in real applications as it can deal with reverberation in time-varying acoustic environments. However, the computational complexity of AWPE is high, which may be a problem for its implementation in real-time systems. This paper presents some new insights into AWPE-based speech dereverberation by introducing the concepts of Kronecker product and partially time-varying filtering. It then develops two algorithms for dereverberation with lower complexity than AWPE. The significant contributions of this work are as follows. First, we propose a Kronecker product filtering framework for speech dereverberation, where the linear prediction filter is formulated as the Kronecker product of two sets of shorter filters. Second, we propose a partially time-varying Kronecker product filter for dereverberation. Instead of estimating the entire linear prediction filter as in the conventional method, the proposed one only needs to update part of the filter. The proposed approaches can significantly reduce the computational complexity without sacrificing dereverberation performance as compared to AWPE. Simulation results validate the theoretical analysis and justify the advantages of the new methods. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | On Differential Beamforming With Nonuniform Linear Microphone ArraysabstractWhile differential beamforming with uniform linear arrays (ULAs) has been widely studied, there is little work so far regarding the design of differential beamformers with nonuniform linear arrays (NULAs). This paper attempts to shed some light on the principles of differential beamforming with NULAs. We define spatial difference operators with NULAs, where any order of the spatial difference of the observation signals can be represented as the product of a nonuniform spatial difference operator matrix and the observation vector. Consequently, the design of differential beamformers is performed in two stages. In the first one, a nonuniform spatial difference operator matrix is applied to the array observations, thereby yielding differential signals. In the second stage, beamformers are designed and applied to the obtained differential signals to optimize the array performance. Based on the defined spatial difference operators, we derive from some performance metrics a family of differential beamformers with NULAs, which include the maximum directivity factor (DF), the maximum white noise gain (WNG), and the maximum front-to-back ratio (FBR) differential beamformers. To compromise between the DF and array robustness, we also derive the parameterized maximum DF and parameterized maximum FBR differential beamformers. The null-constraint maximum DF and WNG differential beamformers are also developed so that some nulls can be placed in specified directions for interference suppression. Simulation results validate the theoretical analysis and justify the properties of the proposed methods. Jilu Jin, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Combined Differential Beamforming With Uniform Linear Microphone ArraysabstractWhile differential beamformers have been widely used in voice communication and human-machine speech interface systems to enhance speech signals of interest, how to design such beamformers that on the one hand can achieve the highest possible directivity factor (DF) and on the other hand are able to obtain a certain level of white noise gain (WNG), so that they are robust enough to sensors’ self noise and array imperfections is still a challenging issue. This paper studies the problem of robust differential beamforming with small-size arrays to achieve a high DF. It presents a method for the design of differential beamformers with uniform linear arrays. We first generate differential pressure signals by applying the recently developed forward spatial difference operator to the outputs of the array with pressure sensors. The pressure microphone observation signals and the differential pressure signals are then put together, and a combined beamformer is subsequently designed, which consists of two subbeamformers, one operates on the pressure microphone observations and the other on the differential pressure signals. A new class of combined differential beamformers are introduced, which can achieve different levels of compromises between DF and WNG using an adjustable parameter. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
ICASSP | 1 |
| 2021 | Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone ArraysabstractDifferential beamformers with concentric circular microphone arrays (CCMAs) are desirable for use in various applications since they can form frequency-invariant spatial responses, have better beam steering flexibility than linear arrays, and suffer less with beampattern irregularity and white noise amplification than circular microphone arrays (CMAs). The methods developed previously for differential beamforming with CCMAs are based on the series expansion. Such methods need to know the analytic form of the target beam-pattern, which may not be accessible in practice. Furthermore, expansion error may lead to erroneous solution, which can cause noise amplification instead of reduction. In this paper, we extend our recently developed beamforming method for CMAs to the design of differential beamformers with CCMAs, which takes advantage of the symmetric null constraints from the beampattern. Simulations are performed to justify the properties of the proposed approach. Xuehan Wang, Gongping Huang, Israel Cohen, Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2021 | On the Design of Square Differential Microphone Arrays with a Multistage StructureabstractThis paper studies the problem of designing square differential microphone arrays (SDMAs). It presents a multistage approach, which first divides an SDMA composed of M2microphones into (M − 1)2subarrays with each subarray being a 2 × 2 square array formed by four adjacent microphones. Then, differential beamforming is performed with each subarray in the first-stage. The first-stage differential beamformers’ outputs are subsequently used as the inputs of the second stage to form (M − 2)2subarrays and a second-stage differential beamforming is then performed. Continuing this process till the (M −1)th stage, we obtain the final output of the SDMA. The SDMA designed in such a multistage structure has two important properties. First, the global weighting matrix is equal to the two dimensional convolution of weighting matrices from the first stage to the last one. Second, the global beampattern is equal to the product of beampatterns from all stages. Consequently, we can combine different kinds of beamformers in different stages and have better control of the performance metrics. Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
ICASSP | 2 |
| 2021 | Time Difference of Arrival Estimation Based on a Kronecker Product DecompositionabstractTime difference of arrival (TDOA) estimation, which often serves as the fundamental step for a source localization or a beamforming system, has a significant practical importance in a wide spectrum of applications. To deal with reverberation, the TDOA estimation problem is often transformed into one of identifying the relative acoustic impulse responses. This letter presents a method to efficiently identify the relative acoustic impulse response between two microphones for TDOA estimation based on the so-called Kronecker product decomposition. By decomposing the relative impulse response into a series of Kronecker products of shorter filters, the original channel identification problem with a long impulse response is converted into one of identifying a number of short filters. Since the TDOA information is embedded only in the direct path of the relative impulse response, the dimension of the Kronecker product decomposition can be very small and, as a result, the developed algorithm is expected to work well in real environments with a small number of data snapshots. Xianrui Wang, Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
IEEE Signal Process. Lett. | 2 |
| 2021 | Robust Dereverberation With Kronecker Product Based Multichannel Linear PredictionabstractReverberation impairs not only the speech quality, but also intelligibility. The weighted-prediction-error (WPE) method, which estimates the late reverberation component based on a multichannel linear predictor, is by far one of the most effective algorithms for dereverberation. Generally, the WPE prediction filter in every short-time-Fourier-transform (STFT) subband has to be long enough to estimate accurately the late reverberation component. As a consequence, WPE is computationally expensive, which makes it difficult to implement into real-time embedded or edge computing devices. Moreover, WPE is sensitive to additive noise and its performance may suffer from dramatic degradation even in environments where the signal-to-noise ratio (SNR) is high. To address these drawbacks, this letter proposes to decompose the multichannel linear prediction filter as a Kronecker product of a temporal (interframe) prediction filter and a spatial filter. An iterative algorithm is then developed to optimize the two filters. In comparison with the original WPE algorithm, the presented method not only exhibits better performance in terms of dereverberation and robustness to additive noise, as there are fewer parameters to estimate for a given number of observation signal samples, but is also computationally more efficient, since the dimensions of the covariance matrices after Kronecker product decomposition are smaller. Wenxing Yang, Gongping Huang, Jingdong Chen, Jacob Benesty, Israel Cohen, Walter Kellermann |
IEEE Signal Process. Lett. | 2 |
| 2021 | On a Particular Family of Differential Beamformers With Cardioid-Like and No-Null PatternsabstractDifferential microphone arrays (DMAs), which are responsive to the differential acoustic pressure fields, have been used in a wide range of applications related to audio and speech. The core part of a DMA is the so-called differential beamformer, which is generally designed by placing a number of nulls in its beampattern to attenuate noise from some directions. But the presence of these nulls may cause some great issues, e.g., leading to suboptimal performance if the interference/noise is incident from directions other than the nulls' directions, and making the beamformer less robust to sensors' self noise and array imperfections. To overcome these problems, this letter is devoted to the design of differential beamformers with no nulls in its beampattern. A design method and its multistage implementation are presented and analyzed. An improved solution is then developed, which is able to form frequency-invariant beampatterns with no nulls in the frequency range of speech signals. Simulations are provided to illustrate the properties of the developed methods. Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2021 | On the Robustness of the Superdirective BeamformerabstractIn microphone array beamforming, a high directional gain is always desired for acoustic noise and reverberation suppression; as a result, the superdirective beamformer has been of great interest in many applications. However, this beamformer is well known to be very sensitive to array imperfections. While much effort has been made to improve its robustness, it is still a major problem. This paper is essentially devoted to the study of the robustness of the superdirective beamformer and derivation of better ways to deal with this important issue. We first prove that any distortionless fixed beamformer can be written as the sum of two orthogonal beamformers, i.e., the sum of the classical delay-and-sum (DS) beamformer and a reduced-rank beamformer. Based on this property, different kinds of robust superdirective beamformers are then developed. We also show that the robust design problem can be transformed into a quadratic eigenvalue problem (QEP), which leads to a solution that achieves the maximum possible directivity factor (DF) while meets the white noise gain (WNG) constraint over a frequency band of interest. Xi Chen 0128, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Steering Study of Linear Differential Microphone ArraysabstractDifferential microphone arrays (DMAs) can achieve high directivity and frequency-invariant spatial response with small apertures; they also have a great potential to be used in a wide spectrum of applications for high-fidelity sound acquisition. Although many efforts have been made to address the design of linear DMAs (LDMAs), most developed methods so far only work for the situation where the source of interest is incident from the endfire direction. This paper studies the steering problem of differential beamformers with linear microphone arrays. We present new insights into beam steering of LDMAs and propose a series of steerable differential beamformers. The major contributions of this paper are as follows. 1) A series of ideal functions are defined to describe the ideal, target beampatterns of LDMAs. 2) We prove that first-order differential beamformers with linear microphone arrays are not steerable and their mainlobes can only be at the endfire directions. 3) We deduce the fundamental conditions for designing steerable differential beamformers with LDMAs. 4) We develop a method to design steerable beamformers with LDMAs using null constraints. Simulations and experiments validate the properties of the developed method. Jilu Jin, Gongping Huang, Xuehan Wang, Jingdong Chen, Jacob Benesty, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Beamforming with Cube Microphone Arrays Via Kronecker Product DecompositionsabstractMicrophone arrays combined with beamforming have been widely used to solve many important acoustic problems in a wide range of applications. Much effort has been devoted in the literature to microphone array beamforming, among which the Kronecker product beamforming method developed recently has demonstrated some interesting properties. Generally, this method decomposes the global beamforming filter into a Kronecker product of a number of sub-beamforming filters, each of which corresponds to a virtual subarray and can be designed individually. This decomposition not only reduces significantly the number of beamforming coefficients, but also can be explored to improve the robustness and flexibility of beamforming. This paper extends Kronecker product beamforming from two-dimensional arrays into three-dimensional cube arrays. We consider two decompositions, i.e., fully and partially separable ones. The former decomposes the entire array into three linear subarrays while the latter decomposes the entire array into a linear subarray and a planar one. Then, for each case, we derive the Kronecker product maximum white noise gain beamformer, the Kronecker product approximate maximum directivity factor (DF) beamformer, the Kronecker product null-steering beamformer, and the Kronecker product iterative maximum DF beamformer. Simulation results demonstrate the properties and advantages of the proposed beamformers. Xuehan Wang, Jacob Benesty, Jingdong Chen, Gongping Huang, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | A New Class of Differential BeamformersabstractDifferential microphone arrays (DMAs) have been used in a wide range of applications for high-fidelity acoustic signal acquisition and enhancement. In the design of differential beamformers, three of the widely used measures are the directivity factor (DF), the front-to-back ratio (FBR), and the white noise gain (WNG). The former two have been used to obtain optimal differential beamformers, e.g., the hypercardioid and supercardioid, and the third one is generally used to analyze and control the robustness of the beamformer with respect to array imperfections due to sensors' self noise, mismatch among sensors, and sensors' placement errors. In this paper, we present a new measure called directivity factor and front-to-back ratio (DFBR), which is a generalization of DF and FBR. With this new measure, three different kinds of beamformers are derived. The first one is the maximum DFBR beamformer, which is deduced by maximizing DFBR with a joint diagonalization method. The second one is the ψ-cardioid beamformer, which is the maximum DFBR beamformer corresponding to a distortionless constraint. The last one is the reduced-rank differential beamformer, which is obtained by properly choosing the dimension of the signal subspace and maximizing WNG subject to the distortionless constraint. The developed beamformers have many interesting properties, which are justified by both simulations and experiments. Wenxing Yang, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Differential Beamforming From the Beampattern Factorization PerspectiveabstractDifferential beamformers have demonstrated a great potential in forming frequency-invariant beampatterns and achieving high directivity factors. Most conventional approaches design differential beamformers in such a way that their beampatterns resemble a desired or target beampattern. In this paper, we show how to design differential beamformers by simply taking advantage of the fact that the beampattern is actually a particular form of an exponential polynomial. Thanks to this quite obvious formulation, a target beampattern is not really needed while the zeros of the exponential polynomial and/or its factorization are fully exploited. The advantage of this factorization is twofold. First, it gives the relation between the beamformer and the roots of the polynomial, so the former can be directly determined from the latter, which seems natural and convenient. Second, based on this factorization, we propose a new formulation of the beamforming filter, which decomposes the filter into shorter ones with the Kronecker product. This formulation is very general, and many well-known beamformers such as the differential and delay-and-sum (DS) ones can be derived from it. Furthermore, the new formulation allows one to combine different kinds of beamformers together, which gives a great flexibility in forming different beampatterns and achieving a compromise among the directivity factor (DF), white noise gain (WNG), and frequency invariance. Jacob Benesty, Jingdong Chen, Gongping Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | On the Design of 3D Steerable Beamformers With Uniform Concentric Circular Microphone ArraysabstractCircular microphone arrays (CMAs) and concentric CMAs (CCMAs) have been used in a wide range of applications such as smartspeakers and teleconferencing systems because of their flexible steering ability. Although many efforts have been devoted to beamforming with CCMAs, most existing methods consider only the 2-dimensional (2D) case and assume that the sound sources of interest are in the same plane as the sensor array (generally the horizontal plane), which often does not hold true in practical applications. This paper deals with the problem of beamforming with uniform CCMAs (UCCMAs) in the 3-dimensional (3D) space to control the steering of the spatial response and meanwhile form frequency-invariant beampatterns for processing broadband acoustic and speech signals. The major contributions of this work are summarized as follows: 1) it presents an analysis based on the spherical harmonics decomposition about the Nth-order optimal and steerable directivity patterns; 2) a beamforming method is developed in which the beamformer's coefficients are identified by solving a linear system of equations formed by approximating the Nth-order optimal target beampattern with the beamformer's beampattern while the resulting beampattern can be steered flexibly in the 3D space; 3) the sufficient and necessary condition on the array geometry and sensors' placement are given to ensure that the beamformer exists and is unique; and 4) the analytical forms of the directivity factor (DF) and white noise gain (WNG) of the resulting beamformer is given and discussion is presented on what conditions irregularities (deep nulls) in WNG and DF may occur. Simulations are provided to illustrate the property of the developed beamforming methods. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Robust and steerable kronecker product differential beamforming With rectangular microphone arraysabstractDifferential microphone arrays (DMAs), a class of welldesigned small-size arrays combined with differential beamforming, are very useful for processing broadband acoustic, audio, and speech signals in a wide range of applications. However, most efforts in the literature so far have been devoted to linear, circular, and spherical arrays. In this paper, we consider rectangular shapes of planar microphone arrays. Instead of adopting the traditional differential beamforming methods developed in the literature, we present a differential beamforming method based on the so-called Kronecker product. We first decompose the entire rectangular array into two virtual rectangular sub-arrays so that the steering vector of the entire array is the Kronecker product of the steering vectors of the two smaller virtual rectangular sub-arrays. We use the first virtual rectangular array, which is much smaller in size than the entire array but well satisfies the basic requirements for differential beamforming, to design a steerable differential beamformer. For the second virtual rectangular array, we can design either the delay-and-sum (DS) beamformer, which helps to improve the robustness of the global differential beamformer, or an adaptive beamformer, which makes the global differential beamformer adaptive. This method has many interesting properties, particularly the designed beamformer is fully steerable, and its robustness and the array gain can be easily controlled. Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
ICASSP | 1 |
| 2020 | An Improved Solution to the Frequency-Invariant Beamforming with Concentric Circular Microphone ArraysabstractFrequency-invariant beamforming with circular microphone arrays (CMAs) has drawn a significant amount of attention for its steering flexibility and high directivity. However, frequency-invariant beam-forming with CMAs often suffers from the so-called null problem, which is caused by the zeros of the Bessel functions; then, concentric CMAs (CCMAs) are used to deal with this problem. While frequency-invariant beamforming with CCMAs can mitigate the null problem, the beampattern is still suffering from distortion due to s-patial aliasing at high frequencies. In this paper, we find that the spatial aliasing problem is caused by higher-order circular harmonics. To deal with this problem, we take the aliasing harmonics into account and approximate the beampattern with a higher truncation order of the Jacobi-Anger expansion than required. Then, the beam-forming filter is determined by minimizing the errors between the desired directivity pattern and the approximated one. Simulation results show that the developed method can mitigate the distortion of the beampattern caused by spatial aliasing. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2020 | Differential Beamforming on GraphsabstractWe study differential beamforming from a graph perspective. The microphone array used for differential beamforming is viewed as a graph, where its sensors correspond to the nodes, the number of microphones corresponds to the order of the graph, and linear spatial difference equations among microphones are related to graph edges. Specifically, for the first-order differential beamforming with an array of M microphones, each pair of adjacent microphones are directly connected, resulting in M - 1 spatial difference equations. On a graph, each of these equations corresponds to a 2-clique. For the second-order differential beamforming, each three adjacent microphones are directly connected, resulting in M - 2 second-order spatial difference equations, and each of these equations corresponds to a 3-clique. In an analogous manner, the differential microphone array for any order-of-differential beamforming can be viewed as a graph. From this perspective, we then derive a class of differential beamformers, including the maximum white noise gain beamformer, the maximum directivity factor one, and optimal compromising beamformers. Simulations are presented to demonstrate the performance of the derived differential beamformers. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | A Simple Theory and New Method of Differential Beamforming With Uniform Linear Microphone ArraysabstractThis article presents a theoretical study of differential beamforming with uniform linear arrays. By defining a forward spatial difference operator, any order of the spatial difference of the observed signals can be represented as a product of a difference operator matrix and the microphone array observations. Consequently, differential beamforming is implemented in two stages, where the first one obtains spatial difference of the observations and the second stage optimizes the beamformer. The major contributions of this article are as follows. First, we propose a new theory of differential beamforming with uniform linear arrays, which shows clearly the connection between the conventional differential beamforming and the null-constrained differential beamforming methods. This provides some new insight into the design of differential beamformers. Second, we deduce some new differential beamformers, where conventional beamforming may be seen as a particular case. Specifically, we derive the maximum white noise gain (MWNG), maximum directivity factor (MDF), parameterized MDF, and parameterized maximum front-to-back ratio differential beamformers. Third, we further extend the idea of how to design optimal differential beamformers by combining both the observed signals and their spatial differences. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Design of Planar Differential Microphone Arrays With Fractional OrdersabstractDifferential microphone arrays (DMAs) often encounter white noise amplification, especially at low frequencies. If the array geometry and the number of microphones are fixed, one can improve the white noise amplification problem by reducing the DMA order. With the existing differential beamforming methods, the DMA order can only be a positive integer number. Consequently, with a specified beampattern (or a kind of beampattern), reducing this order may easily lead to over compensation of the white noise gain (WNG) and too much reduction of the directivity factor (DF), which is not optimal. To deal with this problem, we present in this article a general approach to the design of DMAs with fractional orders. The major contributions of this article include but are not limited to: 1) we first define a directivity pattern that can achieve a continuous compromise between the pattern corresponding to the maximum DMA order and the omnidirectional pattern; 2) by approximating the beamformer's beampattern with the Jacobi-Anger expansion, we present a method to find the proper differential beamforming filter so that its beampattern matches closely the target directivity pattern of fractional orders; and 3) we show how to determine analytically the proper fractional order of the DMA with a given target beampattern when either the value of the DF or WNG is specified, which is useful in practice to achieve the desired beampattern and spatial gain while maintaining the robustness of the DMA system. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone ArraysabstractSmall aperture circular microphone arrays (CMAs) have been widely used in many applications such as teleconferencing, smartspeakers, and robotics. A critical component of such arrays is the differential beamformer, which can achieve relatively high spatial gains with the same beampatterns at most frequencies. Among different differential beamforming approaches that were developed in the literature, the minimum-norm one has attracted much interest as it can deal better with sensors' self noise, sensor mismatch, and beamformer's irregularity at some frequencies due to the zeros of the Bessel functions. In our previous study, we have investigated the performance of the minimum-norm differential beamformer with uniform CMAs (UCMAs) in the 2-dimensional (2D) space where the sound sources and the sensors are assumed to be in the same plane. But in practice, this assumption is generally not true. So, in this paper, we investigate the properties and limitations of the minimum-norm differential beamformer in the 3-dimensional(3D) space. Through theoretical study as well as simulations, we show that the minimum-norm differential beamformer is effective in dealing with the problem of white noise amplification and irregularity of the beampatterns and the directivity factor (DF) if the steering angles are within or near the sensor plane, but it becomes less and less effective as the beamformer is steered away from this plane. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 1 |
| 2019 | Design of Optimal Linear Differential Microphone Arrays Based Array Geometry OptimizationabstractThis paper presents a method to design optimal linear differential microphone arrays (DMAs) by optimizing the array geometry. By constraining the DMA beamformer to achieve a given target value of the directivity factor (DF) with a specified target frequency-invariant beampattern while achieving also the highest possible white noise gain (WNG), an optimization algorithm is developed, which consists of the following two steps. 1) The full frequency band of interest is divided into a few subbands. At every subband, the entire linear array is divided into subarrays and the number of subarrays depends on the total number of the sensors and the order of the DMA. A cost function is then defined, which is minimized to determine what subarray produces the optimal performance. 2) The subband optimal subarrays are then combined across the entire frequency band to form a fullband cost function, from which the geometry of the entire array is optimized. These two steps are repeated with the particle swarm optimization (PSO) algorithm until the desired array performance is reached. Simulation results demonstrate that the proposed method can obtain the target DF with a frequency-invariant beampattern over a wide band of frequencies while maintaining a reasonable level of WNG. Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2019 | On the Design of Flexible Kronecker Product Beamformers with Linear Microphone ArraysabstractThis paper proposes a method for the design of flexible Kronecker product beamformers based on the decomposition of the steering vector of a physical array as a Kronecker product of steering vectors of two smaller virtual arrays. With this decomposition, the global beamforming filter is designed by optimizing the two sub-beamformers in a cascaded manner, which can offer much flexibility to control the performance of beamforming or control the compromise between different, conflicted performance measures. In comparison with a recently developed method that restricts the number of microphones of the given physical array to a multiplication of two integers, each corresponding to the number of sensors of one virtual array, the approach in this work decomposes the physical array in such a way that the sensors in the two virtual arrays may share positions and the number of microphones of the physical array can be any positive integer. Simulations demonstrate the properties of the proposed approach. Wenxing Yang, Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
ICASSP | 2 |
| 2018 | On the Design of Robust Steerable Frequency-Invariant Beampatterns with Concentric Circular Microphone ArraysabstractThis paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs). We develop a beamforming algorithm based on an optimal approximation of the beamformer's beampattern with the Jacobi-Anger expansion. In comparison with the existing frequency-invariant beamformers with either circular microphone arrays (CMAs) or CCMAs, the developed algorithm offers the following advantages: 1) it can mitigate the deep-null problem encountered in CMAs and therefore has a consistent directivity factor over the frequency range of speech signals; 2) it is more flexible in terms of steering flexibility and the resulting beampattern can be steered to any direction; and 3) it does not require the microphones in different rings of the CCMA to be aligned, which is very useful in practice, particularly when microphone arrays with small and compact apertures have to be used. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 1 |
| 2018 | Insights Into Frequency-Invariant Beamforming With Concentric Circular Microphone ArraysabstractThis paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs) and presents an approach to the design of frequency-invariant and symmetric beampatterns. We first apply the Jacobi-Anger expansion to each ring of the CCMA to approximate the beampattern. The beamformer is then designed by using all the expansions from different rings. In comparison with the existing work in the literature where a Jacobi-Anger expansion of the same order is applied to different rings, here in this contribution the order of the Jacobi-Anger expansion at a ring is related to its number of sensors and, as a result, the expansion order at different rings may be different. The developed approach is rather general. It is not only able to mitigate the deep nulls problem in the directivity factor and the white noise gain, that is common to circular microphone arrays (CMAs), and improve the steering flexibility, but is also flexible to use in practice where a smaller ring can have less microphones than a larger one. We discuss the conditions for the design ofNth-order symmetric beampatterns and examples of frequency-invariant beampatterns with commonly used array geometries such as CMAs, CMAs with a sensor at the center, and CCMAs. We show the advantage of adding one microphone at the center of either a CMA or a CCMA, i.e., circumventing the deep nulls problem caused by the 0th-order Bessel function. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Study of the frequency-domain multichannel noise reduction problem with the householder transformationabstractThis paper presents an approach to the multichannel noise reduction problem. It first transforms the multichannel noisy speech signals into the frequency domain. A Householder transformation is then constructed, which converts the multichannel coefficients in each frequency bin into two components: one dominated by speech and the other dominated by noise. A Wiener filter is subsequently formed to achieve an estimate of the noise in the speech dominated component from the noise dominated component. The enhanced speech is then obtained by subtracting the noise estimate from the speech dominated component. This approach consists of two critical steps: construction of the Householder transformation and formation of the noise reduction Wiener filter. If the source incidence angle is known a priori, the Householder transformation can be directly constructed using the steering vector and the optimal estimate of the signal of interest can then be obtained by applying the Wiener filter. If the source incidence angle is not known a priori, the Householder transformation can be constructed from a hypothesized incidence angle. Then, the optimal signal estimate is obtained by searching the maximum of the variance of the enhanced signal with the Wiener filter in the interested range of the incidence angle. Gongping Huang, Jacob Benesty, Jingdong Chen |
ICASSP | 1 |
| 2017 | On the Design of Frequency-Invariant Beampatterns With Uniform Circular Microphone ArraysabstractThis paper deals with two critical issues about uniform circular arrays (UCAs): frequency-invariant response and steering flexibility. It focuses on some optimal design of frequency-invariant beampatterns in any desired direction along the sensor plane. The major contributions are as follows. 1) We explain how to include the steering information in the desired directivity pattern. 2) We show that the optimal approximation of the beamformer's beampattern with a UCA from a least-squares error perspective is the Jacobi-Anger expansion. 3) We develop an approach to the design of any desired symmetric directivity pattern, where the deduced beampattern is almost frequency invariant and its main beam can be pointed to any wanted direction in the sensor plane. 4) With the proposed approach, we derive an explicit form of the white noise gain (WNG) and the directivity factor (DF), and explain clearly the white noise amplification problem at low frequencies and the DF degradation at high frequencies. The analysis also indicates that increasing the number of microphones can always improve the WNG. We show that the proposed method is a generalization of circular differential microphone arrays. The relationship between the proposed method and the so-called circular harmonics beamformers is also discussed. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Subspace superdirective beamformers based on joint diagonalizationabstractAlthough they have been intensively studied and used in many applications due to their high directivity factor (DF), superdirective beamformers are sensitive to sensor noise and mismatch between sensors. This paper studies the problem of superdirective beamforming combined with the joint diagonalization method. We develop a subspace superdirective beamforming approach, which can achieve a good compromise between a high DF and white noise amplification. Simulations are performed to justify our theoretical analysis and demonstrate the good properties of this subspace superdirective beamforming approach. Changlei Li, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 3 |
| 2016 | Superdirective Beamforming Based on the Krylov MatrixabstractSuperdirective beamforming has attracted a significant amount of research interest in speech and audio applications, since it can maximize the directivity factor (DF) given an array geometry and, therefore, is efficient in dealing with signal acquisition in diffuse-like noise environments. However, this beamformer is very sensitive to sensor self-noise and mismatch among sensors, which considerably restricts its use in practical systems. This paper develops an approach to superdirective beamforming based on the Krylov matrix. We show that the columns of a proposed Krylov matrix, which span a chosen dimension of the whole space, are interesting beamformers; consequently, all different linear combinations of those columns lead to beamformers that have good properties. In particular, we develop the Krylov maximum white noise gain and Krylov maximum DF beamformers, which are obtained by maximizing the WNG and the DF, respectively. By properly choosing the dimension of the Krylov subspace, the developed beamformers that can make a compromise between reasonable values of the DF and white noise amplification. We also extend the basic idea to the design of the Krylov maximum front-to-back ratio, parametric superdirective, and parametric supercardioid beamformers. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Investigation of a parametric gain approach to single-channel speech enhancementabstractThis paper investigates a parametric gain approach to single-channel noise reduction in the frequency domain. In comparison with the traditional parametric Wiener gain, the major novelty of this presented approach is that the parametric gain is formulated to estimate the noise by using the mean-squared error (MSE) between the noise and the noise estimate. The enhanced signal is then obtained by subtracting the noise estimate from the noisy observation signal. We show that this new method is more practical to implement and can produce better noise reduction performance as compared to the traditional parametric Wiener filtering techniques if the order of the parametric gain is not equal to 1. If the order is 1, the parametric gain is similar to the traditional Wiener gain. Simulation results are presented to illustrate the properties of this new approach. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 1 |
| 2015 | Optimal single-channel noise reduction filtering matrices from the pearson correlation coefficient perspectiveabstractThis paper studies the problem of single-channel noise reduction in the time domain, where an estimate of a vector of the desired clean speech is achieved by filtering a frame of the noisy signal with a rectangular filtering matrix. The core issue with this problem formulation is then the estimation of the optimal filtering matrix. The squared Pearson correlation coefficient (SPCC) is used. We show that different optimal filtering matrices can be derived by maximizing or minimizing the SPCCs between different signals. For example, maximizing the SPCC between the enhanced signal and the filtered speech gives the reduced-rankWiener and minimum distortion (MD) filtering matrices while minimizing the SPCC gives the minimum noise (MN) and another reduced-rank Wiener filtering matrices. Simulation results are presented to illustrate the properties of these filtering matrices. Jiaolong Yu, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 3 |
| 2014 | Examples of optimal noise reduction filters derived from the squared Pearson correlation coefficientabstractThis paper studies the problem of single-channel noise reduction in the time domain. Based on some orthogonal decomposition developed recently and the squared Pearson correlation coefficient (SPCC), several noise reduction filters are derived. We will show that the optimization of the SPCC leads to the Wiener, minimum variance distortionless response (MVDR), minimum noise (MN), minimum uncorrelated speech and noise (MUSN), and linearly constrained minimum variance (LCMV) filters. We also compare the Wiener and MVDR filters derived from the SPCC to their counterparts derived from the mean-square error (MSE) criterion. Simulations are provided to illustrate the performance of all the deduced noise reduction filters. Jiaolong Yu, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 3 |
| 2014 | A family of maximum SNR filters for noise reductionabstractThis paper is devoted to the study and analysis of the maximum signal-to-noise ratio (SNR) filters for noise reduction both in the time and short-time Fourier transform (STFT) domains with one single microphone and multiple microphones. In the time domain, we show that the maximum SNR filters can significantly increase the SNR but at the expense of tremendous speech distortion. As a consequence, the speech quality improvement, measured by the perceptual evaluation of speech quality (PESQ) algorithm, is marginal if any, regardless of the number of microphones used. In the STFT domain, the maximum SNR filters are formulated by considering the interframe information in every frequency band. It is found that these filters not only improve the SNR, but also improve the speech quality significantly. As the number of input channels increases so is the gain in SNR as well as the speech quality. This demonstrates that the maximum SNR filters, particularly the multichannel ones, in the STFT domain may be of great practical value. Gongping Huang, Jacob Benesty, Tao Long 0004, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2008 | Design of Steerable Linear Differential Microphone Arrays With Omnidirectional and Bidirectional SensorsabstractThis paper is dedicated to the design of fully steerable linear differential microphone arrays (LDMAs). We analyze the steerable ideal spatial responses and explain why conventional LDMAs consisting of only omnidirectional microphones have limited steering ability. In order to circumvent this limitation, we suggest to use both omnidirectional and bidirectional (with a dipole shaped directivity pattern) microphones. We discuss the minimum numbers of omnidirectional and bidirectional sensors required for achieving steerable spatial responses and present a method to design fully steerable differential beamformers with LDMAs through the Jacobi-Anger series expansion. Simulations validate the presented technique and the steering flexibility of the designed LDMAs. Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |