EDBT 2026 Demo / reviewers in the wild / expert
Jingdong Chen
dblp:33/5656
· DBLP profile ↗
238ranked-venue papers
21as first author
122since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 163 · 15 first-author · 85 since 2021Artificial intelligence and machine learning · 108 · 10 first-author · 62 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAN: Self-Calibrated AutoregressioN for High-Quality Visual GenerationabstractHuman artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCAN) model capable of self-evaluating and refining generation quality without regenerating the entire image. We unify image token generation and quality evaluation into a single autoregressive model, formulating both tasks as categorical prediction problems. During inference, the model first generates a coarse initial image, then iteratively refines the lowest-quality patches until satisfactory image quality is achieved. Experimental results demonstrate that SCAN effectively handles diverse real-world generation errors and achieves a promising balance between image quality and speed. For example, SCAN-XL achieves an FID of 2.10 and an IS of 326.1, surpassing the LlamaGen-XL by 1.29 (+38%) in FID and 99.0 (+43.6%) in IS, with a 5.6× speedup (19.76s to 3.56s). Compared to recent works, SCAN improves FID and speed by +18.3% and +23% over VAR-d20, and by +7% and +46% over RandAR-XL. Zhanzhou Feng, Qingpei Guo, Jingdong Chen, Ming Yang 0007, Shiliang Zhang |
AAAI | 3 |
| 2026 | HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMsabstractWhile Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of empathetic, context-aware responses. Here we introduce HumanSense, a comprehensive benchmark designed to evaluate the human-centered perception and interaction capabilities of MLLMs, with a particular focus on deep understanding of extended multimodal contexts and the formulation of rational feedback. Our evaluation reveals that leading MLLMs still have considerable room for improvement, particularly for advanced interaction-oriented tasks. Supplementing visual input with audio and text information yields substantial improvements, and Omni-modal models show advantages on these tasks.Furthermore, grounded in the observation that appropriate feedback stems from a contextual analysis of the interlocutor's needs and emotions, we posit that reasoning ability serves as the key to unlocking it. We devise a multi-stage, modality-progressive reinforcement learning approach, resulting in HumanSense-Omni-Reasoning, which substantially enhances performance on higher-level understanding and interactive tasks. Additionally, we observe that successful reasoning processes appear to exhibit consistent thought patterns. By designing corresponding prompts, we also enhance the performance of non-reasoning models in a training-free manner. Ruobing Zheng, Jingdong Chen, Le Wang 0003 |
AAAI | 6 |
| 2026 | UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and PerceptionabstractThe remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual and textual modalities, especially in scenarios involving complex semantic instructions. However, existing approaches often rely heavily on vision-language models (VLMs) or modular designs for semantic guidance, leading to fragmented architectures and computational inefficiency. To address these challenges, we propose UniAlignment, a unified multimodal generation framework within a single diffusion transformer. UniAlignment introduces a dual-stream diffusion training strategy that incorporates both intrinsic-modal semantic alignment and cross-modal semantic alignment, thereby enhancing the model's cross-modal consistency and instruction-following robustness. Additionally, we present SemGen-Bench, a new benchmark specifically designed to evaluate multimodal semantic consistency under complex textual instructions. Extensive experiments across multiple tasks and benchmarks demonstrate that UniAlignment outperforms existing baselines, underscoring the significant potential of diffusion models in unified multimodal generation. Xinyang Song, Weining Wang 0001, Shaozhen Liu, Jingdong Chen, Qi Li 0005, Zhenan Sun |
AAAI | 6 |
| 2026 | Robust Distributed Cooperative Classification With Learned Compressed-Feature DiffusionabstractCooperative inference in distributed sensor networks is challenged by limited communication bandwidth and the risk of node failures. This paper introduces Compressed Feature Diffusion for Decentralized Classification (CFD-DC), a novel framework that addresses these challenges. Each node performs local inference using its own features and compressed feature representations received from other nodes. Our approach relies on two key components: first, a trainable feature compressor at each node that learns compact representations, reducing communication while preserving critical discriminative information; second, an adaptive node weighting mechanism that dynamically adjusts the influence of local and remote features, providing robustness to unreliable or failed nodes. Experiments on multi-view image classification and a simulated multi-node underwater acoustic target classification task demonstrate the effectiveness of the framework. The results show competitive performance compared to centralized and state-of-the-art multi-view methods, reduced communication costs, and superior robustness in scenarios with node failures. Xiling Yao, Jie Chen 0022, Jingdong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | On adaptive multichannel dereverberation based on dichotomous coordinate descent and data-reuse techniques
Wenxing Yang, Jilu Jin, Jingdong Chen, Jacob Benesty |
Signal Process. | 4 |
| 2026 | Design of Low-Rank differential beamformers with constrained directivity or robustness
Kunlong Zhao, Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
Signal Process. | 5 |
| 2026 | A third-order tensor decomposition based algorithm for speech dereverberation
Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
Signal Process. | 4 |
| 2026 | A Phase-Based Feature for Gas Detection Under Unstable Preheating Condition for E-NosesabstractElectronic noses (e-noses) utilizing metal oxide semiconductor (MOS) gas sensors are widely used; however, they typically require several days of preheating after power-on, whether during initial startup or following a power interruption, to reach a stable operation. Such prolonged preheating significantly limits immediate deployment, particularly in portable systems, and has rarely been systematically addressed in the existing literature. To tackle this challenge, we propose a phase-based feature extraction method to capture stable signal patterns under temperature modulation to address baseline drift during the unstable preheating stage. Building on this feature, we develop a detection method that markedly outperforms conventional magnitude-based techniques, which typically exhibit very low detection probabilities during early preheating. The proposed method achieves high-precision gas detection within just 14 hours, and reaches even better detection probability in 2 hours that conventional methods require 154 hours to match when most commercial sensors remain in the pre-conditioning stage. By incorporating the phase-based feature extraction strategy, the required stabilization time is substantially shortened, enabling rapid and reliable gas detection, and facilitating real-time deployment of e-nose systems in critical applications such as emergency response and industrial safety. Lihua Guo, Zhengqiao Zhao, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2026 | Test-Time Learning for Outlier DetectionabstractIn this work, the concept of test-time learning is presented, wherein Machine-Learning (ML) models are constructed by involving unlabeled test samples. Based on this concept, we propose an unsupervised method called Local Augment (LA) designed to improve the performance of trained outlier detectors at the prediction stage without altering the trained models or accessing the training data. LA operates under the only assumption that the model should produce similar outputs for similar inputs, implying that the prediction of a given sample can be enhanced by the predictions for its similar samples. Specifically, LA boosts outlier detection performance during prediction by fusing the outlier score of a given sample with the scores of synthetically neighboring samples generated by adding random perturbations to the given sample. This simple method demonstrates an average improvement of +0.04 Area Under the Receiver Operating Characteristic curve (AUROC) across 22 real-world datasets for all 11 tested detectors. Notably, this represents the pioneering work of enhancing ML models during the prediction stage without the need to modify the trained models or access the training dataset. This work opens up new possibilities for addressing existing bottleneck problems in various ML tasks beyond outlier detection in diverse domains. Jiawei Yang 0001, Jingdong Chen, Susanto Rahardja |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography EstimationabstractFeature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority of existing methods focus on improving coarse feature representation rather than the fine-matching module. Prior fine-matching techniques, which rely on point-to-patch matching probability expectation or direct regression, often lack precision and do not guarantee the continuity of feature points across sequential images. To address this limitation, this paper concentrates on enhancing the fine-matching module in the semi-dense matching framework. We employ a lightweight and efficient homography estimation network to generate the perspective mapping between patches obtained from coarse matching. This patch-to-patch approach achieves the overall alignment of two patches, resulting in a higher sub-pixel accuracy by incorporating additional constraints. By leveraging the homography estimation between patches, we can achieve a dense matching result with low computational cost. Extensive experiments demonstrate that our method achieves higher accuracy compared to previous semi-dense matchers. Meanwhile, our dense matching results exhibit similar end-point-error accuracy compared to previous dense matchers while maintaining semi-dense efficiency. Xiaolong Wang 0013, Lei Yu 0005, Jiangwei Lao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Yu Zhang 0018, Ming Yang 0007 |
AAAI | 7 |
| 2025 | Reversing Flow for Image RestorationabstractImage restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, which introduces inefficiency and complexity. In this work, we propose ResFlow, a novel image restoration framework that models the degradation process as a deterministic path using continuous normalizing flows. ResFlow augments the degradation process with an auxiliary process that disambiguates the uncertainty in HQ prediction to enable reversible modeling of the degradation process. ResFlow adopts entropy-preserving flow paths and learns the augmented degradation flow by matching the velocity field. ResFlow significantly improves the performance and speed of image restoration, completing the task in fewer than four sampling steps. Extensive experiments demonstrate that ResFlow achieves state-of-the-art results across various image restoration benchmarks, offering a practical and efficient solution for real-world applications. Haina Qin, Wenyang Luo, Jingdong Chen, Ming Yang 0007, Bing Li 0001, Weiming Hu 0004 |
CVPR | 5 |
| 2025 | MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video GenerationabstractThe image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such models on large-scale video set in the wild. Traditional metrics, e.g., SSIM or optical flow, are hard to generalize to arbitrary videos, while, it is very tough for human annotators to label the abstract motion intensity neither. Furthermore, the motion intensity shall reveal both local object motion and global camera movement, which has not been studied before. This paper addresses the challenge with a new motion estimator, capable of measuring the decoupled motion intensities of objects and cameras in video. We leverage the contrastive learning on randomly paired videos and distinguish the video with greater motion intensity. Such a paradigm is friendly for annotation and easy to scale up to achieve stable performance on motion estimation. We then present a new I2V model, named MotionStone, developed with the decoupled motion estimator. Experimental results demonstrate the stability of the proposed motion estimator and the state-of-the-art performance of MotionStone on I2V generation. These advantages warrant the decoupled motion estimator to serve as a general plug-in enhancer for both data processing and video generation training. Shuwei Shi, Biao Gong, Zizheng Yang, Yuyuan Li 0001, Jingwen He, Kecheng Zheng, Jingdong Chen, Ming Yang 0007, Yinqiang Zheng |
CVPR | 10 |
| 2025 | Mimir: Improving Video Diffusion Models for Precise Text UnderstandingabstractText serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features from text encoders yet struggle with limited text comprehension. The recent success of large language models (LLMs) showcases the power of decoder-only transformers, which offers three clear benefits for text-to-video (T2V) generation, namely, precise text understanding resulting from the superior scalability, imagination beyond the input text enabled by next token prediction, and flexibility to prioritize user interests through instruction tuning. Nevertheless, the feature distribution gap emerging from the two different text modeling paradigms hinders the direct use of LLMs in established T2V models. This work addresses this challenge with Mimir, an end-to-end training framework featuring a carefully tailored token fuser to harmonize the outputs from text encoders and LLMs. Such a design allows the T2V model to fully leverage learned video priors while capitalizing on the text-related capability of LLMs. Extensive quantitative and qualitative results demonstrate the effectiveness of Mimir in generating high-quality videos with excellent text comprehension, especially when processing short captions and managing shifting motions. Project page: https://lucaria-academy.github.io/Mimir/ Biao Gong, Yutong Feng, Kecheng Zheng, Shuwei Shi, Yujun Shen, Jingdong Chen, Ming Yang 0007 |
CVPR | 8 |
| 2025 | SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingabstractOpen-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two challenges. 1) Existing RS semantic categories are limited, particularly for pixel-level interpretation datasets. 2) Distinguishing among diverse RS spatial regions solely by language space is challenging due to the dense and intricate spatial distribution in open-world RS imagery. To address the first issue, we develop a fine-grained RS interpretation dataset, Sky-SA, which contains 183,375 high-quality local image-text pairs with full-pixel manual annotations, covering 1,763 category labels, exhibiting richer semantics and higher density than previous datasets. Afterwards, to solve the second issue, we introduce the vision-centric principle for vision-language modeling. Specifically, in the pre-training stage, the visual self-supervised paradigm is incorporated into image-text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across open-category texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarthOV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https://github.com/zqcrafts/SkySense-O. Qi Zhu 0010, Jiangwei Lao, Deyi Ji, Lixiang Ru, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Dong Liu 0002, Feng Zhao 0004 |
CVPR | 9 |
| 2025 | Exploring Spectral Signatures of Chinese liquor using Machine Learning and SHapley Additive exPlanationsabstractChinese liquor holds great cultural and economic significance globally. The accurate classification of aroma types and alcohol content is crucial for quality control in Chinese liquor production. To address limitations such as subjectivity and sensor drift in current methods, this study introduces a noninvasive, efficient, and objective approach using Near-Infrared Hyperspectral Imaging (NIR-HSI) to identify alcohol content and aroma types in Chinese liquor. Specifically, we create a comprehensive NIR hyperspectral dataset of Chinese liquor samples and train various machine learning algorithms to classify liquor samples based on their spectral features. Results show that the XGBoost model achieves the optimal performance in most of the experiments. The spectral signatures of various Chinese liquors are explored using SHapley Additive exPlanations (SHAP). We find that spectral bands, such as the band in the range of 1140-1160 nm, make important contributions to the aroma type classification, which are associated with C-H and CO bond stretching vibrations, providing valuable insights into the molecular basis of Chinese liquor analysis. Our dataset is publicly available at https://github.com/wangyunjeff/Chinese-Liquor-NIR-HSI-Dataset. Danlei Chen, Linruize Tang, Zhengqiao Zhao, Jingdong Chen |
ICASSP | 6 |
| 2025 | Advances in Microphone Array Processing and Multichannel Speech EnhancementabstractThis paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It provides valuable insights into the progression and future direction of these areas. The paper examines foundational developments in microphone array design and optimization, showcasing innovations that improved sound acquisition and enhanced speech intelligibility in noisy and reverberant environments. It then introduces recent advancements and cutting-edge research in the field, particularly the integration of deep learning techniques such as all-neural beamformers. The paper also explores critical applications, discussing their evolution and current state-of-the-art technologies that significantly impact user experience. Finally, the paper outlines future research directions, identifying challenges and potential solutions that could drive further innovation in these fields. By providing a comprehensive overview and forward-looking perspective, this paper aims to inspire ongoing research and contribute to the sustained growth and development of microphone arrays and multichannel speech enhancement. Gongping Huang, Jesper Rindom Jensen, Jingdong Chen, Jacob Benesty, Mads Græsbøll Christensen, Akihiko Sugiyama, Gary W. Elko, Tomas Gänsler |
ICASSP | 3 |
| 2025 | Microphone Array Beamforming for Speech Enhancement Based on Dynamic Mode DecompositionabstractMicrophone array beamforming is widely used to extract desired speech signals from noisy environments. While most research in this area focuses on utilizing spatial information, less attention is given to the intrinsic physical mechanisms underlying microphone array observations. This paper aims to address this gap by exploring these underlying factors through dynamic mode decomposition (DMD). Our contributions are twofold. 1) We develop a DMD-based signal model for microphone arrays to capture the relationships between observation signals at adjacent microphones. 2) We introduce a DMD-based preprocessing method and a corresponding beamforming approach based on this model. Simulation results show that our proposed method significantly enhances performance compared to conventional beamforming techniques. Wei Liu 0177, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2025 | On the Design of Low-Rank Differential Beamformers with Nonuniform Linear Microphone ArraysabstractKronecker product beamforming is an effective technique for designing beamformers with nonuniform linear arrays (NULAs). However, current techniques are restricted to NULAs with specific configurations, where the steering vector of the array is represented as a Kronecker product of steering vectors from smaller virtual arrays. This paper overcomes these constraints by proposing a novel approach to designing Kronecker product beamformers for NULAs from a low-rank perspective. Our approach involves decomposing the NULA into overlapping subarrays and organizing the sensor signals from these subarrays into a matrix. We then apply filters to both sides of this matrix to produce an output, which is then converted into a low-rank beamforming process. This method is highly adaptable and can be utilized for NULAs with any number of microphones. Hanchen Pei, Gongping Huang, Jilu Jin, Jacob Benesty, Jingdong Chen |
ICASSP | 6 |
| 2025 | Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech EnhancementabstractSuperdirective beamformers are highly effective at suppressing directional interference and diffuse noise, but their practical use is often constrained by the problem of white noise amplification. Robust superdirective beamforming methods typically address this by imposing a constraint on the white noise gain (WNG). However, determining the appropriate WNG threshold in varying noise environments remains unclear. This paper introduces a data-driven approach to estimating the optimal WNG threshold. Subsequently, a more versatile and robust superdirective beamformer is developed by solving a quadratic eigenvalue problem (QEP). Experimental results show that this method outperforms traditional superdirective beamformers, which rely on a WNG threshold set through a fixed search range. Importantly, this approach functions as a distortionless beamformer, maintaining high fidelity of the desired acoustic signal and allowing for additional post-filtering if required. Hanchen Pei, Gongping Huang, Jilu Jin, Zhizheng Wu 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 6 |
| 2025 | Design and Optimization of Superdirective Beamforming and Post-Filtering for Speech EnhancementabstractSuperdirective beamformers, used with small microphone arrays, are highly attractive due to their high directivity and frequency-invariant beampatterns, making them well-suited for processing broadband acoustic and speech signals. However, these beamformers are very sensitive to array imperfections such as sensor mismatches and self-noise. To improve robustness, robust superdirective (RSD) beamformers have been developed, employing techniques such as diagonal loading or white-noise-gain constraints during their derivation. Although RSD beamformers offer enhanced robustness compared to classical superdirective beamformers, they cannot achieve the maximum directivity factor and lose some frequency-invariant properties, resulting in a beamwidth that is wider at low frequencies and narrower at high frequencies. As a result, RSD beamformers do not fully meet the criteria of true superdirective beamformers, providing less effective noise reduction and introducing some speech distortion. Post-filtering methods have been developed to improve noise reduction after RSD beamforming, but they often fail to address the distortion issues, especially when the speech source deviates from the array’s look direction. To overcome this limitation, this paper proposes a joint optimization approach that combines post-filtering with RSD beamformers. By using the output of RSD beamformers as input data and considering various deviations in look directions and array mismatches, we train a post-filtering network to further enhance the beamformer’s output. Experimental results on speech enhancement demonstrate the effectiveness and robustness of the proposed method. Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2025 | DOA Estimation Based on Enhanced SRP-MVDR Using Kronecker Product Decomposition for Large Rectangular Microphone ArraysabstractDirection-of-arrival (DOA) estimation is a key process in microphone array systems. The steered response power-based minimum variance distortionless response (SRP-MVDR) method performs very well in challenging acoustic environments but suffers from exponential complexity as the number of microphones increases. To improve the efficiency of SRP-MVDR for real-time applications, we propose a Kronecker product-based SRP-MVDR (SRP-KPMVDR) method designed for large rectangular microphone arrays. This approach begins with a rank-one approximation that represents the signal covariance matrix of a rectangular microphone array in Kronecker product form, which is essential for SRP-MVDR estimation. By utilizing the Kronecker product properties, the complex matrix inversion in SRP-MVDR is simplified to the inversion of two smaller matrices, significantly reducing computational complexity. Simulation results show that the SRP-KPMVDR method achieves comparable performance to the traditional SRP-MVDR while greatly decreasing the computational demands. Yichen Zeng, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2025 | Radiation and Directivity Analysis of a Vibrating Dome-Shaped Radiator Mounted on an Infinite BaffleabstractAccurate modeling and analysis of a radiator mounted on an infinite baffle are crucial for understanding its acoustic radiation characteristics. This paper investigates the radiation behavior of a convex dome-shaped radiator in such a condition, showing that, under the far-field approximation, the pressure field is the three-dimensional Fourier transform of the axisymmetric surface velocity distribution. The study includes a comparison of various typical velocity distributions in terms of their directivity factor (DF) and radiated sound power. Current velocity distributions often suffer from nulls in the DF, so we provide a detailed analysis to uncover the causes of this issue. To address this, we propose an equalization filter designed to smooth the DF across the entire frequency range. Simulations are performed to validate the theoretical findings and to showcase the improved performance of the proposed approach. Junqing Zhang, Wen Zhang 0002, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2025 | Design of Robust Differential Beamformers with Microphone Arrays of Arbitrary Planar GeometryabstractDifferential microphone arrays (DMAs) have garnered significant attention in recent research and development due to their high directivity and frequency-invariant beampatterns. However, DMAs frequently encounter substantial white noise amplification, which limits their practical applications. This paper addresses this issue by introducing a general method for designing robust DMAs with microphone arrays of arbitrary planar topology. The proposed approach approximates the beampattern using the Jacobi-Anger series expansion and constrains the white noise gain (WNG) to a specified value. This minimizes the error between the beampattern and the ideal directivity pattern while ensuring a reasonable level of robustness. A closed-form solution for the robust differential beamformer filter is derived using the quadratic eigenvalue problem (QEP) method. Simulation results demonstrate the feasibility and effectiveness of the proposed approach. Kunlong Zhao, Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2025 | CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidanceabstractSemi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we propose a novel pipeline, CasP, which leverages cascaded correspondence priors for guidance. Specifically, the matching stage is decomposed into two progressive phases, bridged by a region-based selective cross-attention mechanism designed to enhance feature discriminability. In the second phase, one-to-one matches are determined by restricting the search range to the one-to-many prior areas identified in the first phase. Additionally, this pipeline benefits from incorporating high-level features, which helps reduce the computational costs of low-level feature extraction. The acceleration gains of CasP increase with higher resolution, and our lite model achieves a speedup of $\sim2.2\times$ at a resolution of 1152 compared to the most efficient method, ELoFTR. Furthermore, extensive experiments demonstrate its superiority in geometric estimation, particularly with impressive cross-domain generalization. These advantages highlight its potential for latency-sensitive and high-robustness applications, such as SLAM and UAV systems. Code is available at https://github.com/pq-chen/CasP. Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yingying Pei, Xinyi Liu 0002, Yongxiang Yao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002 |
ICCV | 10 |
| 2025 | When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
Xue Yang 0005, Qi Zhu 0010, Jingdong Chen, Yansheng Li 0001 |
ICCV | 7 |
| 2025 | SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingabstractThe multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for each data modality, leading to redundancy and inefficient parameter utilization. Moreover, prevalent pre-training methods typically apply self-supervised learning (SSL) techniques from natural images without adequately accommodating the characteristics of remote sensing (RS) images, such as the complicated semantic distribution within a single RS image. In this work, we present SkySense V2, a unified MM-RSFM that employs a single transformer backbone to handle multiple modalities. This backbone is pre-trained with a novel SSL strategy tailored to the distinct traits of RS data. In particular, SkySense V2 incorporates an innovative adaptive patch merging module and learnable modality prompt tokens to address challenges related to varying resolutions and limited feature diversity across modalities. In additional, we incorporate the mixture of experts (MoE) module to further enhance the performance of the foundation model. SkySense V2 demonstrates impressive generalization abilities through an extensive evaluation involving 16 datasets over 7 tasks, outperforming SkySense by an average of 1.8 points. Lixiang Ru, Lei Yu 0005, Yansheng Li 0001, Jingdong Chen |
ICCV | 7 |
| 2025 | Animate-X: Universal Character Image Animation with Enhanced Motion RepresentationabstractCharacter image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used in industries like gaming and entertainment. Our in-depth analysis suggests to attribute this limitation to their insufficient modeling of motion, which is unable to comprehend the movement pattern of the driving video, thus imposing a pose sequence rigidly onto the target character. To this end, this paper proposes $\texttt{Animate-X}$, a universal animation framework based on LDM for various character types (collectively named $\texttt{X}$), including anthropomorphic characters. To enhance motion representation, we introduce the Pose Indicator, which captures comprehensive motion pattern from the driving video through both implicit and explicit manner. The former leverages CLIP visual features of a driving video to extract its gist of motion, like the overall movement pattern and temporal relations among motions, while the latter strengthens the generalization of LDM by simulating possible inputs in advance that may arise during inference. Moreover, we introduce a new Animated Anthropomorphic Benchmark ($\texttt{$A^2$Bench}$) to evaluate the performance of $\texttt{Animate-X}$ on universal and widely applicable animation images. Extensive experiments demonstrate the superiority and effectiveness of $\texttt{Animate-X}$ compared to state-of-the-art methods. Biao Gong, Xiang Wang 0012, Shiwei Zhang 0001, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, Ming Yang 0007 |
ICLR | 8 |
| 2025 | The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Shilong Wu, Hang Chen 0001, Jun Du 0002, Chin-Hui Lee 0001, Shinji Watanabe 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg |
INTERSPEECH | 7 |
| 2025 | On the Design of a Robust Superdirective Beamformer and Topology Parameter Optimization with Frustum-Shaped Microphone Arrays Featuring Multiple Rings
Kunlong Zhao, Gongping Huang, Jingdong Chen, Jacob Benesty, Zoran Cvetkovic |
INTERSPEECH | 4 |
| 2025 | Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
Ruobing Zheng, Jingdong Chen, Ming Yang 0007 |
ACM Multimedia | 4 |
| 2025 | VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional AnnotationsabstractVideo aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-based methods. This study introduces VADB, the largest video aesthetic database with 10,490 diverse videos annotated by 37 professionals across multiple aesthetic dimensions, including overall and attribute-specific aesthetic scores, rich language comments and objective tags. We propose VADB-Net, a dual-modal pre-training framework with a two-stage training strategy, which outperforms existing video quality assessment models in scoring tasks and supports downstream video aesthetic assessment tasks. The dataset and source code are available at https://github.com/BestiVictory/VADB. Qianqian Qiao, Yihang Bo, Bao Peng, Heng Huang 0002, Longteng Jiang, Huaye Wang, Jingdong Chen, Xin Jin 0015 |
NeurIPS | 8 |
| 2025 | ARGenSeg: Image Segmentation with Autoregressive Image Generation ModelabstractWe propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework.
Prior works integrating image segmentation into multimodal large language models (MLLMs) typically employ either boundary points representation or dedicated segmentation heads.
These methods rely on discrete representations or semantic prompts fed into task-specific decoders, which limits the ability of the MLLM to capture fine-grained visual details.
To address these challenges,
we introduce a segmentation framework for MLLM based on image generation, which naturally produces dense masks for target objects.
We leverage MLLM to output visual tokens and detokenize them into images using an universal VQ-VAE,
making the segmentation fully dependent on the pixel-level understanding of the MLLM.
To reduce inference latency,
we employ a next-scale-prediction strategy to generate required visual tokens in parallel.
Extensive experiments demonstrate that our method surpasses prior state-of-the-art approaches on multiple segmentation datasets with a remarkable boost in inference speed, while maintaining strong understanding capabilities. Xiaolong Wang 0013, Lixiang Ru, Kaixiang Ji, Jingdong Chen, Jun Zhou 0011 |
NeurIPS | 6 |
| 2025 | VideoMAR: Autoregressive Video Generation with Continuous TokensabstractMasked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored.
Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored.
In this paper, we propose \textbf{VideoMAR}, a concise and efficient decoder-only autoregressive image-to-video model with continuous tokens, composing temporal frame-by-frame and spatial masked generation.
We first identify temporal causality and spatial bi-directionality as the first principle of video AR models, and propose the next-frame diffusion loss for the integration of mask and video generation.
Besides, the huge cost and difficulty of long sequence autoregressive modeling is a basic but crucial issue. To this end, we propose the temporal short-to-long curriculum learning and spatial progressive resolution training, and employ progressive temperature strategy at inference time to mitigate the accumulation error.
Furthermore, VideoMAR replicates several unique capacities of language models to video generation.
It inherently bears high efficiency due to simultaneous temporal-wise KV cache and spatial-wise parallel generation, and presents the capacity of spatial and temporal extrapolation via 3D rotary embeddings.
On the VBench-I2V benchmark, VideoMAR surpasses the previous state-of-the-art (Cosmos I2V) while requiring significantly fewer parameters ($9.3\%$), training data ($0.5\%$), and GPU resources ($0.2\%$). Hu Yu 0001, Biao Gong, Hangjie Yuan, Weilong Chai, Jingdong Chen, Kecheng Zheng, Feng Zhao 0004 |
NeurIPS | 6 |
| 2025 | PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
Peiyuan Zhang, Xue Yang 0005, Yi Yu 0010, Qingyun Li, Yue Zhou 0005, Xiaosong Jia, Jingdong Chen, Xiang Li 0041, Junchi Yan, Yansheng Li 0001 |
Int. J. Comput. Vis. | 9 |
| 2025 | A MISO acoustic echo cancellation algorithm based on a two-layer filter decomposition
Zi Cao, Tianci Yan, Jingdong Chen, Jacob Benesty |
Signal Process. | 4 |
| 2025 | An update rule for multiple source variances estimation using microphone arrays
Fan Zhang 0001, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
Speech Commun. | 3 |
| 2025 | MPPCAD: Minimum Power Pattern Constrained Adaptive Differential BeamformingabstractThis paper investigates the design of adaptive differential beamforming using small-spacing linear microphone arrays. We express the differential beamformer as a linear function of the target beampattern coefficients through orthogonal polynomial expansions. Consequently, the design of the beamformer reduces to optimizing these coefficients. To ensure that the maximum array response consistently aligns with the look direction, we derive constraints on the target beampattern coefficients, resulting in two convex sets for the first two orders of beampatterns. This approach uncovers numerous effective beampatterns beyond traditional options such as dipole, cardioid, supercardioid, and hypercardioid. By minimizing the power of the array output while adhering to the beampattern constraints, we develop the MPPCAD beamformer. Simulation results demonstrate that the proposed beamformer significantly enhances speech quality compared to classical differential beamformers. Fan Zhang 0001, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2025 | Non-Intrusive Speech Quality Assessment Based on Deep Neural Networks for Speech CommunicationabstractTraditionally, speech quality evaluation relies on subjective assessments or intrusive methods that require reference signals or additional equipment. However, over recent years, non-intrusive speech quality assessment has emerged as a promising alternative, capturing much attention from researchers and industry professionals. This article presents a deep learning-based method that exploits large-scale intrusive simulated data to improve the accuracy and generalization of non-intrusive methods. The major contributions of this article are as follows. First, it presents a data simulation method, which generates degraded speech signals and labels their speech quality with the perceptual objective listening quality assessment (POLQA). The generated data is proven to be useful for pretraining the deep learning models. Second, it proposes to apply an adversarial speaker classifier to reduce the impact of speaker-dependent information on speech quality evaluation. Third, an autoencoder-based deep learning scheme is proposed following the principle of representation learning and adversarial training (AT) methods, which is able to transfer the knowledge learned from a large amount of simulated speech data labeled by POLQA. With the help of discriminative representations extracted from the autoencoder, the prediction model can be trained well on a relatively small amount of speech data labeled through subjective listening tests. Fourth, an end-to-end speech quality evaluation neural network is developed, which takes magnitude and phase spectral features as its inputs. This phase-aware model is more accurate than the model using only the magnitude spectral features. A large number of experiments are carried out with three datasets: one simulated with labels obtained using POLQA and two recorded with labels obtained using subjective listening tests. The results show that the presented phase-aware method improves the performance of the baseline model and the proposed model with latent representations extracted from the adversarial autoencoder (AAE) outperforms the state-of-the-art objective quality assessment methods, reducing the root mean square error (RMSE) by 10.5% and 12.2% on the Beijing Institute of Technology (BIT) dataset and Tencent Corpus, respectively. The code and supplementary materials are available at https://github.com/liushenme/AAE-SQA. Miao Liu 0007, Jing Wang 0037, Fei Wang 0030, Jingdong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Towards Better Vision-Inspired Vision-Language ModelsabstractVision-language (VL) models have achieved unprece-dented success recently, in which the connection module is the key to bridge the modality gap. Nevertheless, the abun-dant visual clues are not sufficiently exploited in most existing methods. On the vision side, most existing approaches only use the last feature of the vision tower, without using the low-level features. On the language side, most existing meth-ods only introduce shallow vision-language interactions. In this paper, we present a vision-inspired vision-language con-nection module, dubbed as VIVL, which efficiently exploits the vision cue for VL models. To take advantage of the lower-level information from the vision tower, a feature pyramid extractor (FPE) is introduced to combine features from differ-ent intermediate layers, which enriches the visual cue with negligible parameters and computation overhead. To en-hance VL interactions, we propose deep vision-conditioned prompts (DVCP) that allows deep interactions of vision and language features efficiently. Our VIVL exceeds the previous state-of-the-art method by 18.1 CIDEr when training from scratch on the COCO caption task, which greatly improves the data efficiency. When used as a plug-in module, VIVL consistently improves the performance for various backbones and VL frameworks, delivering new state-of-the-art results on multiple benchmarks, e.g., NoCaps and VQAv2. Yun-Hao Cao, Kaixiang Ji, Chuanyang Zheng, Jiajia Liu 0002, Jian Wang 0108, Jingdong Chen, Ming Yang 0007 |
CVPR | 7 |
| 2024 | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryabstractPrior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primar-ily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pretrained on a curated multimodal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multimodal spatiotemporal encoder taking temporal sequences of opti-cal and Synthetic Aperture Radar (SAR) data as input. This encoder is pretrained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multimodal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thor-ough evaluation encompassing 16 datasets over 7 tasks, from single- to multimodal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pretrained weights to facilitate future research and Earth Observation applications. Xin Guo 0010, Jiangwei Lao, Bo Dang 0002, Lei Yu 0005, Lixiang Ru, Liheng Zhong, Dingxiang Hu, Huimei He, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002, Yansheng Li 0001 |
CVPR | 13 |
| 2024 | Learning Dynamic Tetrahedra for High-Quality Talking Head SynthesisabstractRecent works in implicit representations, such as Neural Radiance Fields (NeRF), have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters, since the lack of explicit geometric constraints poses a fundamental challenge in accurately modeling complex facial deformations. In this paper, we introduce Dynamic Tetrahedra (DynTet), a novel hybrid representation that encodes explicit dynamic meshes by neural networks to ensure geometric consistency across various motions and viewpoints. DynTet is parameterized by the coordinate-based networks which learn signed distance, deformation, and material texture, anchoring the training data into a predefined tetrahedra grid. Leveraging Marching Tetrahedra, DynTet efficiently decodes textured meshes with a consistent topology, enabling fast rendering through a differentiable rasterizer and supervision via a pixel loss. To enhance training efficiency, we incorporate classical 3D Morphable Models to facilitate geometry learning and define a canonical space for simplifying texture learning. These advantages are readily achievable owing to the effective geometric representation employed in DynTet. Compared with prior works, DynTet demonstrates significant improvements in fidelity, lip synchronization, and real-time performance according to various metrics. Beyond producing stable and visually appealing synthesis videos, our method also outputs the dynamic meshes which is promising to enable many emerging applications. Code is available at https://github.com/zhangzc21/DynTet. Ruobing Zheng, Bonan Li, Congying Han, Tiande Guo, Jingdong Chen, Ziwen Liu 0001, Ming Yang 0007 |
CVPR | 8 |
| 2024 | EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yongjun Zhang 0002, Jian Wang 0108, Liheng Zhong, Jingdong Chen, Ming Yang 0007 |
ECCV (68) | 7 |
| 2024 | StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
Wen Li 0024, Muyuan Fang, Biao Gong, Ruobing Zheng, Jingdong Chen, Ming Yang 0007 |
ECCV (28) | 7 |
| 2024 | POA: Pre-training Once for Models of All Sizes
Xin Guo 0010, Jiangwei Lao, Lei Yu 0005, Lixiang Ru, Jian Wang 0108, Guo Ye, Huimei He, Jingdong Chen, Ming Yang 0007 |
ECCV (3) | 9 |
| 2024 | Beamforming Through Online Convex Combination of Differential BeamformersabstractThanks to their high directivity, compact size, and reliable performance, differential microphone arrays (DMAs) have attracted great interest from both industry and academia as they have demonstrated great potential to be used in a wide range of applications for high-fidelity speech acquisition. Nevertheless, in many real-world applications, DMAs powered with fixed differential beamformers are often inadequate in suppressing interference, particularly in environments with multiple or moving sources. To address this issue, this work develops an adaptive convex combination (ACC)-based method, which combines multiple differential beamformers in an online manner for enhanced performance. While the major contribution is a new real-time processing algorithm that facilitates optimal linear combinations of different differential beamformers, making them adapted to dynamic environments, the presented method also provides valuable insights as how to combine different beamformers for online robust implementation. Jilu Jin, Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2024 | A Computationally Efficient Semi-Blind Source Separation Approach for Nonlinear Echo Cancellation Based on an Element-Wise Iterative Source SteeringabstractWhile the semi-blind source separation-based acoustic echo cancellation (SBSS-AEC) has received much research attention due to its promising performance during double-talk compared to the traditional adaptive algorithms, it suffers from system latency and nonlinear distortions. To circumvent these drawbacks, the recently developed ideas on convolutive transfer function (CTF) approximation and nonlinear expansion have been used in the iterative projection (IP)-based semi-blind source separation (SBSS) algorithm. However, because of the introduction of CTF approximation and nonlinear expansion, this algorithm becomes computationally very expensive, which makes it difficult to implement in embedded systems. Thus, we attempt in this paper to improve this IP-based algorithm, thereby developing an element-wise iterative source steering (EISS) algorithm. In comparison with the IP-based SBSS algorithm, the proposed algorithm is computationally much more efficient, especially when the nonlinear expansion order is high and the length of the CTF filter is long. Meanwhile, its AEC performance is as good as that of IP-based SBSS algorithm. Kunxing Lu, Xianrui Wang, Tetsuya Ueda, Shoji Makino, Jingdong Chen |
ICASSP | 5 |
| 2024 | On the Design of Planar Differential Microphone Arrays with Specified Beamwidth or Sidelobe LevelabstractThis paper investigates the problem of designing differential beam-formers with planar microphone arrays to achieve not only the desired target directivity pattern but also control the beamwidth (BW) or sidelobe level (SLL). We first discuss the target directivity patterns and express the Dolph-Chebyshev polynomial based form of target directivity patterns into linear combination of cylindrical harmonics. We then address the problem of designing differential beamformers through beampattern approximation based on the Jacobi-Anger series expansion. Two methods are subsequently developed: the first one involves designing beamformers to achieve the target directivity pattern while minimizing SLL under a pre-specified value of BW and the second one aims to attain the target directivity pattern while minimizing the null-to-null BW under a pre-specified level of SLL. Simulations are carried out to validate the method and the results demonstrate the properties of the proposed method. Xueqin Luo, Jilu Jin, Gongping Huang, Yingke Zhao, Jingdong Chen, Jacob Benesty |
ICASSP | 5 |
| 2024 | A Steered Response Power Approach with Bilinear Prediction-Based Trade-Off Prewhitening for Speaker LocalizationabstractThis paper studies the problem of acoustic source localization in room environments. It presents an improved steered response power (SRP) approach with low-complexity and trade-off prewhitening. This method consists of two steps. In the first one, the linear predictor that is used to model the speech signals is formulated as a bilinear form, and a group of convex-constrained linear prediction sub-models with respect to dual sub-predictors are established to pre-filter microphone signals. The pre-filtered (prewhitened) microphone signals are subsequently used in SRP for speaker localization. Simulation results demonstrate the properties of the presented method: it is robust to reverberation and noise, and is computationally efficient thanks to the bilinear form. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 3 |
| 2024 | The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker ExtractionabstractPrevious Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompted a shift in focus towards the Audio-Visual Target Speaker Extraction (AVTSE) task for the MISP 2023 challenge in ICASSP 2024 Signal Processing Grand Challenges. Unlike existing audio-visual speech enhancement challenges primarily focused on simulation data, the MISP 2023 challenge uniquely explores how front-end speech processing, combined with visual clues, impacts back-end tasks in real-world scenarios. This pioneering effort aims to set the first benchmark for the AVTSE task, offering fresh insights into enhancing the accuracy of back-end speech recognition systems through AVTSE in challenging and real acoustic environments. This paper delivers a thorough overview of the task setting, dataset, and baseline system of the MISP 2023 challenge. It also includes an in-depth analysis of the challenges participants may encounter. The experimental results highlight the demanding nature of this task, and we look forward to the innovative solutions participants will bring forward. Shilong Wu, Hang Chen 0001, Yusheng Dai, Chenyue Zhang, Ruoyu Wang 0029, Hongbo Lan, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Sabato Marco Siniscalchi, Odette Scharenborg, Zhongqiu Wang 0001, Jianqing Gao |
ICASSP | 10 |
| 2024 | Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split NetworkabstractStereophonic music source separation (MSS) is a problem of extracting individual source tracks, e.g. bass, drums, vocals, from a stereo music recording. Deep neural network (DNN) based MSS systems have demonstrated great promise though spatial panning cues and time-frequency spectral structures in stereo music have not yet been fully explored in such systems and methods. This paper presents a spatially-informed MSS method using a bridging band-split neural network that incorporates both spatial and spectral information. The spatial panning angles of each target source are used as input of the network, along with the time-frequency spectrograms. Moreover, the inter-track correlations are exploited for further performance improvement. Experiments show that the proposed method outperforms significantly the baseline systems as the result of using spatial cues, spectral characteristics, and inter-track relationships. Yichen Yang 0010, Xianrui Wang, Wen Zhang 0002, Shoji Makino, Jingdong Chen |
ICASSP | 6 |
| 2024 | Directional Gain Based Noise Covariance Matrix Estimation for MVDR BeamformingabstractThis paper is devoted to the problem of noise covariance matrix (NCM) estimation. It proposes a time-frequency masking based approach. We first present an optimal mask function based on the mean-squared error criterion. To estimate this mask, we employ the recently developed directional gain method based on the knowledge of the signal incident angle. To demonstrate the effectiveness of the proposed NCM estimator, we integrate it into the minimum variance distortionless response (MVDR) beamformer. The speech enhancement results in noise-plus-interference environments show the advantages of the proposed method over two baseline beamforming algorithms. Fan Zhang 0001, Chao Pan 0001, Jacob Benesty, Jingdong Chen |
ICASSP | 4 |
| 2024 | Differential Beamforming with Null Constraints for Spherical Microphone ArraysabstractDifferential microphone arrays (DMAs) can measure both the acoustic pressure field and the differential acoustic pressure fields, which gives them great advantages in a wide range of applications for acoustic and speech signal acquisition. The core component of DMAs is the so-called differential beamformer, the design of which typically involves taking into account the a priori knowledge about the array geometry and the desired directivity pattern that is related to the differential sound field to respond. This paper deals with the design of differential beamformers with spherical microphone arrays. It presents a novel design approach based on the null constraints formed from the desired directivity pattern. In comparison with the exiting methods, the proposed approach only requires the information of the zeros in the beampattern, which provides notable flexibility and convenience for spherical DMA design in practical applications. Xueqin Luo, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 4 |
| 2024 | LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic ConstraintsabstractIntegrating first-order logic constraints (FOLCs) with neural networks is a crucial but challenging problem since it involves modeling intricate correlations to satisfy the constraints. This paper proposes a novel neural layer, LogicMP, which performs mean-field variational inference over a Markov Logic Network (MLN). It can be plugged into any off-the-shelf neural network to encode FOLCs while retaining modularity and efficiency. By exploiting the structure and symmetries in MLNs, we theoretically demonstrate that our well-designed, efficient mean-field iterations greatly mitigate the difficulty of MLN inference, reducing the inference from sequential calculation to a series of parallel tensor operations. Empirical results in three kinds of tasks over images, graphs, and text show that LogicMP outperforms advanced competitors in both performance and efficiency. Weidi Xu, Lele Xie, Jianshan He, Hongting Zhou, Taifeng Wang, Xiaopei Wan, Jingdong Chen, Chao Qu |
ICLR | 8 |
| 2024 | Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual RecognitionabstractLong-tailed recognition (LTR) aims to learn balanced models from extremely unbalanced training data. Fine-tuning pretrained foundation models has recently emerged as a promising research direction for LTR. However, we observe that the fine-tuning process tends to degrade the intrinsic representation capability of pretrained models and lead to model bias towards certain classes, thereby hindering the overall recognition performance. To unleash the intrinsic representation capability of pretrained foundation models, in this work, we propose a new Parameter-Efficient Complementary Expert Learning (PECEL) for LTR. Specifically, PECEL consists of multiple experts, where individual experts are trained via Parameter-Efficient Fine-Tuning (PEFT) and encouraged to learn different expertise on complementary sub-categories via the proposed sample-aware logit adjustment loss. By aggregating the predictions of different experts, PECEL effectively achieves a balanced performance on long-tailed classes. Nevertheless, learning multiple experts generally introduces extra trainable parameters. To ensure parameter efficiency, we further propose a parameter sharing strategy which decomposes and shares the parameters in each expert. Extensive experiments on 4 LTR benchmarks show that the proposed PECEL can effectively learn multiple complementary experts without increasing the trainable parameters and achieve new state-of-the-art performance. Lixiang Ru, Xin Guo 0010, Lei Yu 0005, Jiangwei Lao, Jian Wang 0108, Jingdong Chen, Yansheng Li 0001, Ming Yang 0007 |
ACM Multimedia | 7 |
| 2024 | Accelerating Pre-training of Multimodal LLMs via Chain-of-SightabstractThis paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs).
Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales.
This architecture not only leverages global and local visual contexts effectively, but also facilitates the flexible extension of visual tokens through a compound token scaling strategy, allowing up to a 16x increase in the token count post pre-training.
Consequently, Chain-of-Sight requires significantly fewer visual tokens in the pre-training phase compared to the fine-tuning phase.
This intentional reduction of visual tokens during pre-training notably accelerates the pre-training process, cutting down the wall-clock training time by $\sim$73\%.
Empirical results on a series of vision-language benchmarks reveal that the pre-train acceleration through Chain-of-Sight is achieved without sacrificing performance, matching or surpassing the standard pipeline of utilizing all visual tokens throughout the entire training process.
Further scaling up the number of visual tokens for pre-training leads to stronger performances, competitive to existing approaches in a series of benchmarks. Kaixiang Ji, Biao Gong, Zhiwu Qing, Kecheng Zheng, Jian Wang 0108, Jingdong Chen, Ming Yang 0007 |
NeurIPS | 8 |
| 2024 | Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerabstractAbstract Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performance of self-attention mechanism in the language field, transformers tailored for visual data have drawn significant attention and triumphed over CNNs in various vision tasks. These vision transformers heavily rely on large-scale pre-training to achieve competitive accuracy, which not only hinders the freedom of architectural design in downstream tasks like object detection, but also causes learning bias and domain mismatch in the fine-tuning stages. To this end, we aim to get rid of the “pre-train and fine-tune” paradigm of vision transformer and train transformer based object detector from scratch. Some earlier works in the CNNs era have successfully trained CNNs based detectors without pre-training, unfortunately, their findings do not generalize well when the backbone is switched from CNNs to a vision transformer. Instead of proposing a specific vision transformer based detector, in this work, our goal is to reveal the insights of training vision transformer based detectors from scratch. In particular, we expect those insights to help other researchers and practitioners, and inspire more interesting research in other fields, such as remote sensing, visual-linguistic pre-training, etc. One of the key findings is that both architectural changes and more epochs play critical roles in training vision transformer based detectors from scratch. Experiments on the MS COCO dataset demonstrate that vision transformer based detectors trained from scratch can also achieve similar performance to their counterparts with ImageNet pre-training. Weixiang Hong 0001, Wang Ren, Jiangwei Lao, Lele Xie, Liheng Zhong, Jian Wang 0108, Jingdong Chen, Honghai Liu 0001 |
Int. J. Comput. Vis. | 7 |
| 2024 | On intrusive speech quality measures and a global SNR based metric
Chao Pan 0001, Jingdong Chen, Jacob Benesty |
Speech Commun. | 2 |
| 2024 | A Closed-Form DOA Estimator Using Spherical Microphone Arrays in the Presence of InterferenceabstractDirection-of-arrival (DOA) estimation is challenging in complex acoustic environments with background noise and interference. Utilizing spherical microphone arrays, closed-form estimators can be derived, which are attractive for practical applications due to their computational efficiency, eliminating the need for exhaustive extremum searching. However, current closed-form estimators are susceptible to interference. To address this issue, we propose an estimator that directly computes the DOA of the desired source using the covariance matrix of the observation signals. This approach effectively mitigates the impact of interference when the covariance matrix is accurately estimated. Simulation results demonstrate the superior performance of the proposed method compared to the subspace pseudo-intensity vector (SSPIV) and relative harmonic coefficients (RHC) methods. Yilong Lu, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2024 | On the Design of Robust Differential Beamformers From the Beampattern Error PerspectiveabstractDifferential microphone arrays (DMAs), which enhance acoustic signals of interest by measuring both the acoustic pressure field and its spatial derivatives, find extensive use in various practical systems and acoustic products. A critical element of DMAs is the differential beamformer, traditionally designed to ensure that the designed beampattern closely matches the desired target directivity pattern. However, such beamformers may lack sufficient robustness in practice. To address the balance between robustness and beampattern accuracy, this letter proposes two types of beamformers: one prioritizes maximizing the white noise gain (WNG) while maintaining a specified mean-squared beampattern error (MSBE), and the other aims to minimize MSBE while adhering to a specified level of WNG. By transforming these design challenges into quadratic eigenvalue problems (QEPs), we derive explicit solutions for the proposed beamformers. Simulations are conducted to illustrate the performance characteristics of these beamformers. Jingli Xie, Junqing Zhang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 5 |
| 2024 | Multi-Source DOA Estimation Using Higher-Order Pseudo Intensity Vector on a Spherical Microphone ArrayabstractDirection-of-arrival (DOA) estimation in environments with multiple sources and strong reverberation remains a great challenge. In this letter, we present a novel feature, the higher-order pseudo-intensity vector (HOPIV), derived from recordings obtained with a spherical microphone array. By exploiting the unique properties of the reactive intensity vector, which is derived from the HOPIV, we present a method to identify time-frequency points that are dominated by the direct path. We then propose a DOA estimation method that leverages the HOPIV's high spatial resolution for improving DOA estimation performance. Simulations and experiments show that the proposed method is able to yield superior performance compared to state-of-the-art techniques, even in highly reverberant environments. Wen Zhang 0002, Jingdong Chen, Mengyao Zhu 0003, Chunjian Li |
IEEE Signal Process. Lett. | 3 |
| 2024 | Smoothed Frame-Level SINR and Its Estimation for Sensor Selection in Distributed Acoustic Sensor NetworksabstractDistributed acoustic sensor network (DASN) refers to a sound acquisition system that consists of a collection of microphones randomly distributed across a wide acoustic area. Theory and methods for DASN are gaining increasing attention as the associated technologies can be used in a broad range of applications to solve challenging problems. However, unlike traditional microphone arrays or centralized systems, properly exploiting the redundancy among different channels in DASN is facing many challenges including but not limited to variations in pre-amplification gains, clocks, sensors' response, and signal-to-interference-plus-noise ratios (SINRs). Selecting appropriate sensors relevant to the task at hand is therefore crucial in DASN. In this work, we propose a speaker-dependent smoothed frame-level SINR estimation method for sensor selection in multi-speaker scenarios, specifically addressing source movement within DASN. Additionally, we devise an approach for similarity measurement to generate dynamic speaker embeddings resilient to variations in reference speech levels. Furthermore, we introduce a novel loss function that integrates classification and ordinal regression within a unified framework. Extensive simulations are performed and the results demonstrate the efficacy of the proposed method in accurately estimating smoothed frame-level SINR dynamically, yielding state-of-the-art performance. Shanzheng Guan, Mou Wang, Zhongxin Bai, Jianyu Wang 0007, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Design of Fully Steerable Differential Beamformers With Linear SuperarraysabstractLinear differential microphone arrays (LDMAs) are commonly integrated into thin and portable devices to achieve high-fidelity speech acquisition. Traditional LDMAs typically consist of only omnidirectional microphones, which impose limitations on their ability to produce steerable spatial responses due to constraints in array element directivity and linear array geometry. A recent solution to this limitation involves integrating both omnidirectional and bidirectional microphones in LDMA design, enabling the creation of steerable spatial responses. This paper extends the core idea of integrating omnidirectional and bidirectional microphones, and develops a more general and comprehensive theory and method for designing steerable LDMAs. It makes two main contributions. Firstly, it introduces a general approach to designing steerable LDMAs, in which any type of directional microphones can be used. Secondly, it gives the minimum number of omnidirectional and directional microphones required to achieve a specific order of steerable LDMA. Simulations validate the proposed method and illustrate how omnidirectional and directional sensors can be combined to form the desired LDMAs. Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | On Semi-Blind Source Separation-Based Approaches to Nonlinear Echo Cancellation Based on Bilinear Alternating OptimizationabstractAcoustic echo cancellation (AEC) is a crucial task in full duplex communications. As conventional linear filtering approaches are ineffective to deal with double-talk, various semi-blind source separation (SBSS)-based AEC algorithms are deceived, most of which are formulated and implemented in the frequency domain based on the multiplicative transfer function (MTF) model for computational efficiency. To avoid large latency and in order to deal with loudspeaker nonlinearities, the convolutive transfer function (CTF) model and odd power series expansion are leveraged, which are employed by numerous SBSS-based nonlinear AEC (SBSS-NAEC) algorithms. Conventional SBSS-NAEC methods estimate the series expansion coefficients and the CTF filter simultaneously making the number of free parameters to estimate large. Hence, the corresponding algorithms are computationally expensive and are difficult to optimize. In this work, we propose to decouple the series expansion coefficients and the CTF filters into a bilinear form and present a bilinear alternating optimization framework for estimating the model parameters. An alternating iterative projection (AIP) algorithm and an alternating element-wise iterative source steering (AEISS) algorithm are proposed. As the bilinear representation consists of less parameters compared to the conventional methods, the proposed algorithms not only improve the AEC performance but also reduce the computational complexity, which is validated by comprehensive simulations and experiments. Xianrui Wang, Yichen Yang 0010, Andreas Brendel, Tetsuya Ueda, Shoji Makino, Jacob Benesty, Walter Kellermann, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2024 | Interference-Controlled Maximum Noise Reduction Beamformer Based on Deep-Learned Interference ManifoldabstractBeamforming has been used in a wide range of applications to extract the signal of interest from microphone array observations, which consist of not only the signal of interest, but also noise, interference, and reverberation. The recently proposed interference-controlled maximum noise reduction (ICMR) beamformer provides a flexible way to control the specified amount of the interference attenuation and noise suppression; but it requires accurate estimation of the manifold vector of the interference sources, which is challenging to achieve in real-world applications. To address this issue, we introduce an interference-controlled maximum noise reduction network (ICMRNet) in this study, which is a deep neural network (DNN)-based method for manifold vector estimation. With densely connected modified conformer blocks and the end-to-end training strategy, the interference manifold is learned directly from the observation signals. This approach, akin to ICMR, adeptly adapts to time-varying interference and demonstrates superior convergence rate and extraction efficacy as compared to the linearly constrained minimum variance (LCMV)-based neural beamformers when appropriate attenuation factors are selected. Moreover, via learning-based extraction, ICMRNet effectively suppresses reverberation components within the target signal. Comparative analysis against baseline methods validates the efficacy of the proposed method. Yichen Yang 0010, Ningning Pan, Wen Zhang 0002, Chao Pan 0001, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic SegmentationabstractIn order to tackle video semantic segmentation task at a lower cost, e.g., only one frame annotated per video, lots of efforts have been devoted to investigate the utilization of those unlabeled frames by either assigning pseudo labels or performing feature enhancement. In this work, we propose a novel feature enhancement network to simultaneously model short- and long-term temporal correlation. Compared with existing work that only leverage short-term correspondence, the long-term temporal correlation obtained from distant frames can effectively expand the temporal perception field and provide richer contextual prior. More importantly, modeling adjacent and distant frames together can alleviate the risk of over-fitting, hence produce high-quality feature representation for the distant unlabeled frames in training set and unseen videos in testing set. To this end, we term our method SSLTM, short for Simultaneously Short- and Long-Term Temporal Modeling. In the setting of only one frame annotated per video, SSLTM significantly outperforms the state-of-the-art methods by 2% ∼ 3% mIoU on the challenging VSPW dataset. Furthermore, when working with a pseudo label based method such as MeanTeacher, our final model only exhibits 0.13% mIoU less than the ceiling performance (i.e., all frames are manually annotated). Jiangwei Lao, Weixiang Hong 0001, Xin Guo 0010, Jian Wang 0108, Jingdong Chen |
CVPR | 6 |
| 2023 | Summary on the Multimodal Information Based Speech Processing (MISP) 2022 ChallengeabstractThe Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visual diarization and recognition (AVDR). The training material was based on previous MISP 2021 recordings, but we have accurately synchronized audio and visual data. Additionally, a new evaluation set was provided. This paper gives an overview of the challenge setup, presents the results, and summarizes the effective techniques employed by the participants. We also analyze the current technical challenges and suggest directions for future research in AVSD and AVDR. Hang Chen 0001, Shilong Wu, Yusheng Dai, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 7 |
| 2023 | A Frequency-Domain Recursive Least-Squares Adaptive Filtering Algorithm Based On A Kronecker Product DecompositionabstractThis paper proposes a frequency-domain recursive least-squares (RLS) adaptive filtering algorithm for identifying time-varying acoustic systems in noisy environments. The Kronecker product (KP) is employed to decompose the model filter of the acoustic channel impulse response into two sets of short sub-filters, based on which a generalized frequency-domain signal model and the associated cost function are established. A KP based RLS algorithm is subsequently deduced. In comparison with the conventional frequency-domain RLS adaptive filter, the presented algorithm is not only computationally more efficient, but also has a faster convergence rate for the identification of acoustic systems regardless of whether the excitation is a white sequence or a speech signal. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 2 |
| 2023 | Switching Kronecker Product Linear Filtering for Multispeaker Adaptive Speech DereverberationabstractDereverberation, a process to mitigate or eliminate the reverberation effect, plays an important role in hands-free speech communication and human-machine interfaces. Tremendous efforts have been devoted to this problem and various methods have been developed over the last three decades. Those methods generally assume that there is only a single speaker in the acoustic environment and, consequently, they suffer from significant performance degradation if multiple speakers participate in the conversation. How to deal with reverberation in multiple-speaker scenarios is still a challenging problem, which is studied in this work. We present a switching multichannel linear prediction filtering method, which designs multiple linear filters with each tracking one speaker. When some speaker is active, the corresponding filter and the weighted cross-correlation matrix are updated while the other filters are kept unchanged. To further improve the performance and reduce complexity, we apply the Kronecker product to decompose every linear prediction filter into a Kronecker product of two shorter filters: one is time-invariant and the other is time-varying. The former is estimated with a batch method (using only a few seconds of speech signal when the corresponding speaker starts to talk in the entire conversation) while a recursive least-squares algorithm is derived for identifying the time-varying set of Kronecker filters. Gongping Huang, Jacob Benesty, Israel Cohen, Emil Winebrand, Jingdong Chen, Walter Kellermann |
ICASSP | 5 |
| 2023 | Spatially Informed Independent vector analysis for Source Extraction based on the convolutive Transfer Function ModelabstractSpatial information can help improve source separation performance. Numerous spatially informed source extraction methods based on the independent vector analysis (IVA) have been developed, which can achieve reasonably good performance in non- or weakly reverberant environments. However, the performance of those methods degrades quickly as the reverberation increases. The underlying reason is that those methods are derived based on the multiplicative transfer function model with a rank-1 assumption, which does not hold true if reverberation is strong. To circumvent this issue, this paper proposes to use the convolutive transfer function (CTF) model to improve the source extraction performance and develop a spatially informed IVA algorithm. Simulations demonstrate the efficacy of the developed method even in highly reverberant environments. Xianrui Wang, Andreas Brendel, Gongping Huang, Yichen Yang 0010, Walter Kellermann, Jingdong Chen |
ICASSP | 6 |
| 2023 | On Multiple-Input/Binaural-Output Antiphasic Speaker Signal ExtractionabstractThis paper studies the problem of target speaker signal exaction and antiphasic rendering with an array of microphones in the scenarios where there are two active speakers. Based on the important findings achieved in the psychoacoustic field as well as our recent works on single-channel speech enhancement, we present a rendering based approach in which a temporal convolutional network (TCN) is trained to take the multiple signals observed by the microphone array as its inputs and generate two output (binaural) signals. The TCN is trained in such a way that, when binaural output signals are listened by the listener with headsets, the speech signal from the desired speaker is perceived on one side of and close to the listener’s head, while the competing speech signal is perceived on the opposite side and also away from the listener’s head. Benefited from rendering and the signal-to-interference ratio (SIR) improvement, this antiphasic binaural presentation enables the listener to better focus on the target speaker’s signal while ignoring the impact of the competing speech. The modified rhyme tests (MRTs) are performed to validate the superiority of the proposed method. Xianrui Wang, Ningning Pan, Jacob Benesty, Jingdong Chen |
ICASSP | 4 |
| 2023 | The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And RecognitionabstractThe Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two tracks: 1) audio-visual speaker diarization (AVSD), aiming to solve "who spoken when" using both audio and visual data; 2) a novel audio-visual diarization and recognition (AVDR) task that focuses on addressing "who spoken what when" with audio-visual speaker diarization results. Both tracks focus on the Chinese language, and use far-field audio and video in real home-tv scenarios: 2-6 people communicating each other with TV noise in the background. This paper introduces the dataset, track settings, and baselines of the MISP2022 challenge. Our analyses of experiments and examples indicate the good performance of AVDR baseline system, and the potential difficulties in this challenge due to, e.g., the far-field video quality, the presence of TV noise in the background, and the indistinguishable speakers. Shilong Wu, Hang Chen 0001, Maokui He, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 7 |
| 2023 | Uncertainty-guided Learning for Improving Image Manipulation DetectionabstractImage manipulation detection (IMD) is of vital importance as faking images and spreading misinformation can be malicious and harm our daily life. IMD is the core technique to solve these issues and poses challenges in two main aspects: (1) Data Uncertainty, i.e., the manipulated artifacts are often hard for humans to discern and lead to noisy labels, which may disturb model training; (2) Model Uncertainty, i.e., the same object may hold different categories (tampered or not) due to manipulation operations, which could potentially confuse the model training and result in unreliable outcomes. Previous works mainly focus on solving the model uncertainty issue by designing meticulous features and networks, however, the data uncertainty problem is rarely considered. In this paper, we address both problems by introducing an uncertainty-guided learning framework, which measures data and model uncertainties by a novel Uncertainty Estimation Network (UEN). UEN is trained under dynamic supervision, and outputs estimated uncertainty maps to refine manipulation detection results, which significantly alleviates the learning difficulties. To our knowledge, this is the first work to embed uncertainty modeling into IMD. Extensive experiments on various datasets demonstrate state-of-the-art performance, validating the effectiveness and generalizability of our method. Kaixiang Ji, Feng Chen 0047, Xin Guo 0010, Yadong Xu, Jian Wang 0108, Jingdong Chen |
ICCV | 6 |
| 2023 | Wall-to-Wall Above-Ground Biomass Estimation with Alos-2 Palsar-2 L-Band SAR Data and GEDIabstractUnder the impact of climate change, monitoring forest carbon stock becomes an important task to evaluate the changes in carbon sequestrated from the atmosphere. Forest carbon stock estimation is still a challenging task, due to limited data sources that have a high correlation with above-ground biomass. With the help of the NASA Global Ecosystem Dynamics Investigation (GEDI) mission, above-ground biomass (AGB) can be measured by using the LiDAR data provided. However, GEDI data is sparse since it only samples about 4% of the Earth’s land surface between 51.6° N&S. Previous studies demonstrated L-Band SAR’s promising ability in retrieving forest stem volumes and estimating above-ground biomass. In this work, we propose a Deep Learning based workflow which utilizes PALSAR-2 L-Band images and GEDI to generate wall-to-wall above-ground biomass maps of North America. The workflow uses Convolutional Neural Network as the DL model and leverages both PALSAR-2 L-Band images and GEDI Relative Heights data to estimate the dense above-ground biomass maps. The results show that, by fusing GEDI Level 2 Relative Heights data with PALSAR-2 L-Band SAR data, it is possible to achieve a significantly high correlation with GEDI level 4 AGB data, as the final R-squared score of our model is as high as 0.83. Xin Guo 0010, Liheng Zhong, Jian Wang 0108, Jingdong Chen |
IGARSS | 5 |
| 2023 | Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NERabstractThe challenge posed by multimodal named entity recognition (MNER) is mainly two-fold: (1) bridging the semantic gap between text and image and (2) matching the entity with its associated object in image. Existing methods fail to capture the implicit entity-object relations, due to the lack of corresponding annotation. In this paper, we propose a bidirectional generative alignment method named BGA-MNER to tackle these issues. Our BGA-MNER consists of image2text and text2image generation with respect to entity-salient content in two modalities. It jointly optimizes the bidirectional reconstruction objectives, leading to aligning the implicit entity-object relations under such direct and powerful constraints. Furthermore, image-text pairs usually contain unmatched components which are noisy for generation. A stage-refined context sampler is proposed to extract the matched cross-modal content for generation. Extensive experiments on two benchmarks demonstrate that our method achieves state-of-the-art performance without image input during inference. Feng Chen 0047, Jiajia Liu 0002, Kaixiang Ji, Wang Ren, Jian Wang 0108, Jingdong Chen |
ACM Multimedia | 6 |
| 2023 | Fine-grained Pseudo Labels for Scene Text RecognitionabstractPseudo-Labeling based semi-supervised learning has shown promising advantages in Scene Text Recognition (STR). Most of them usually use a pre-trained model to generate sequence-level pseudo labels for text images and then re-train the model. Recently, conducting Pseudo-Labeling in a teacher-student framework (a student model is supervised by the pseudo labels from a teacher model) has become increasingly popular, which trains in an end-to-end manner and yields outstanding performance in semi-supervised learning. However, applying this framework directly to Pseudo-Labeling STR exhibits unstable convergence, as generating pseudo labels at the coarse-grained sequence-level leads to inefficient utilization of unlabelled data. Furthermore, the inherent domain shift between labeled and unlabeled data results in low quality of derived pseudo labels. To mitigate the above issues, we propose a novel Cross-domain Pseudo-Labeling (CPL) approach for scene text recognition, which makes better utilization of unlabeled data at the character-level and provides more accurate pseudo labels. Specifically, our proposed Pseudo-Labeled Curriculum Learning dynamically adjusts the thresholds for different character classes according to the model's learning status. Moreover, an Adaptive Distribution Regularizer is employed to bridge the domain gap and improve the quality of pseudo labels. Extensive experiments show that CPL boosts those representative STR models to achieve state-of-the-art results on six challenging STR benchmarks. Besides, it can be effectively generalized to handwritten text. Xiaoxue Chen, Zuming Huang, Lele Xie, Jingdong Chen, Ming Yang 0007 |
ACM Multimedia | 5 |
| 2023 | A binaural heterophasic adaptive beamformer and its deep learning assisted implementation
Jilu Jin, Ningning Pan, Jingdong Chen, Jacob Benesty, Yiqian Yang |
Pattern Recognit. Lett. | 3 |
| 2023 | Dimensionality Reduction of Room Acoustic Impulse Responses and Applications to System IdentificationabstractA room Acoustic Impulse Response (RAIR), which represents the sound propagation channel via direct and reflection paths from a source position to a microphone, plays a leading role in a broad range of acoustic signal processing applications, e.g., echo cancellation. In practical acoustic environments, it is not uncommon that an RAIR may consist of hundreds or even thousands of coefficients, making it challenging to identify and handle. This paper investigates the RAIR dimensionality reduction problem inspired from the concepts of dynamic mode decomposition. The objective is to find effective lower-dimensional representations of RAIRs, which are easier and more robust to identify and equalize. There are two main contributions of this work. First, we present an RAIR dimensionality reduction method. Second, we show how to apply this technique to the problem of acoustic system identification. Simulation results demonstrate that the proposed method is able to improve significantly the performance of acoustic system identification. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2023 | Differential Beamforming From a Geometric PerspectiveabstractDifferential microphone arrays (DMAs) have demonstrated a great potential for solving the high-fidelity sound acquisition problem in a wide range of applications as they possess many good properties such as frequency-independent beampatterns with high directivity. A significant number of efforts have been devoted to the design of DMAs and the associated beamformers. As a result, many different types of DMAs and differential beamforming methods have been developed over the last few decades, some of which have been successfully deployed in real systems and commercial products. However, given an application, how to design a DMA to achieve optimal performances is still an open issue. This work studies the problem of designing linear DMAs (LDMAs) from a geometric perspective. Based on the fundamental observation that most practical and interesting DMA beampatterns have nulls in some directions, we define a criterion based on the orthogonality between the beamforming filter and the steering vector in the nulls' directions. We then derive a family of differential beamformers by optimizing the defined criterion, some of which are well known but derived from a different perspective, while others are new. Simulations and experiments are carried out, and the results validate the proposed method and developed differential beamformers. Jilu Jin, Jacob Benesty, Jingdong Chen, Gongping Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Design of Maximum Directivity Beamformers With Linear Acoustic Vector Sensor ArraysabstractThis paper studies the design of maximum directivity factor (MDF) beamformers based on uniform linear arrays (ULAs) consisting of acoustic vector sensors (AVSs). We first derive the main lobe constraints, which ensure that the beamformer's beampattern achieves a maximum in the look direction, and prove that any beamformer that satisfies the proposed constraints can be written as the sum of two orthogonal beamformers: the maximum white noise gain (MWNG) beamformer and a reduced-rank beamformer. Then, we derive the MDF beamformer by maximizing the directivity factor (DF) under the deduced constraints. We also derive a robust version of the MDF beamformer, which can keep the WNG above a pre-specified level. Compared to the conventional MDF beamformer based on ULAs with omnidirectional microphones, the designed MDF beamformer with uniform linear AVS arrays (ULAVSAs) can steer the beampattern to any look direction in the 3-dimensional space and achieves a higher directivity. The proposed MDF beamformer also outperforms the two-step MDF beamformer with ULAVSAs since it maximizes the DF. The proposed methods are validated through simulations as well as real experiments. Xueqin Luo, Gongping Huang, Jilu Jin, Jingdong Chen, Jacob Benesty, Wen Zhang 0002, Mengyao Zhu 0003, Chunjian Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | CGMM-Based Sound Zone Generation Using Robust Pressure Matching With ATF Perturbation ConstraintsabstractPersonal sound zone (PSZ) refers to the technique that uses an array of loudspeakers and digital signal processing tools to achieve spatial soundfield control. To generate the target sound zones, this technique generally requires to know the acoustic transfer functions (ATFs) between the loudspeakers and the spots where soundfields are to be controlled. In practical applications, however, the true ATFs are never accessible and they have to be measured or estimated. Due to many sophisticated reasons, the measured ATFs generally deviate from the true ones, which may lead to significant degradation in performance of sound zone reproduction. In this work, a robust pressure matching (RPM) algorithm is presented for sound zone generation. It exploits a complex Gaussian mixture model (CGMM) to model the ATFs and their perturbations. The CGMM parameters are estimated using the expectation-maximization (EM) algorithm. To improve the robustness of the pressure matching method, an uncertainty constraint is applied to the ATF estimates and the pressure matching problem is then formulated as one of biconvex optimization. The coordinate descent algorithm is subsequently used to solve the optimization problem, thereby obtaining the optimal control filter. In comparison with the existing pressure matching methods without considering the effect of ATF perturbations, the presented algorithm is able to achieve lower normalized signal distortion energy and higher signal to interference ratio. Numerical simulations justify the effectiveness of the presented algorithm as well as its advantages over the traditional methods. Junqing Zhang, Liming Shi, Mads Græsbøll Christensen, Wen Zhang 0002, Lijun Zhang 0004, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Design of 2D and 3D Differential Microphone Arrays With a Multistage FrameworkabstractDifferential microphone arrays (DMAs) have demonstrated a great potential for high-fidelity acoustic and speech signal acquisition in a wide range of applications since such arrays are able to achieve frequency-invariant beampatterns with high directivity. Consequently, a great number of efforts have been devoted to the design of DMAs and the associated beamformers in the literature. However, most of the methods only work for arrays with particular topologies, e.g., linear, circular, concentric circular, and spherical ones. How to design general two-dimensional (2D) and three-dimensional (3D) DMAs that can measure the desired differential sound field and form the desired spatial response in the 3D space remains an unsolved problem. This paper investigates this problem and presents a multistage design approach. The major contributions of this work are as follows. First, we reexamine the differentials of the acoustic pressure field in the 3D space and derive the general expression of the directivity patterns resulting from the spatial differential operation, which serves as the foundation for differential beamforming with 2D or 3D microphone arrays. Second, we present a multistage approach to the design of 2D and 3D DMAs, and deduce the relationship between the global beamformer and the beamformers at different stages as well as the relationship between their beampatterns. Third, several algorithms are presented for the design of differential as well as robust differential beamformers in this multistage framework. Simulation results validate the proposed approach and justifies its properties. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating DeepfakesabstractMalicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models, leading them to generate distorted outputs. Despite achieving impressive results, these adversarial watermarks have low image-level and model-level transferability, meaning that they can protect only one facial image from one specific deepfake model. To address these issues, we propose a novel solution that can generate a Cross-Model Universal Adversarial Watermark (CMUA-Watermark), protecting a large number of facial images from multiple deepfake models. Specifically, we begin by proposing a cross-model universal attack pipeline that attacks multiple deepfake models iteratively. Then, we design a two-level perturbation fusion strategy to alleviate the conflict between the adversarial watermarks generated by different facial images and models. Moreover, we address the key problem in cross-model optimization with a heuristic approach to automatically find the suitable attack step sizes for different models, further weakening the model-level conflict. Finally, we introduce a more reasonable and comprehensive evaluation method to fully test the proposed method and compare it with existing ones. Extensive experimental results demonstrate that the proposed CMUA-Watermark can effectively distort the fake facial images generated by multiple deepfake models while achieving a better performance than existing methods. Our code is available at https://github.com/VDIGPKU/CMUA-Watermark. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Jingdong Chen, Weisi Lin, Kai-Kuang Ma |
AAAI | 8 |
| 2022 | Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerabstractModeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performances of self-attention mech-anism in the language field, transformers tailored for visual data have drawn numerous attention and triumphed CNNs in various vision tasks. These vision transformers heavily rely on large-scale pre-training to achieve competitive accuracy, which not only hinders the freedom of architectural design in downstream tasks like object detection, but also causes learning bias and domain mismatch in the fine-tuning stages. To this end, we aim to get rid of the “pre-train & fine-tune” paradigm of vision transformer and train transformer based object detector from scratch. Some earlier work in the CNNs era have successfully trained CNNs based detectors without pre-training, unfortunately, their findings do not generalize well when the backbone is switched from CNNs to vision transformer. Instead of proposing a specific vision transformer based detector, in this work, our goal is to reveal the insights of training vision transformer based detectors from scratch. In particular, we expect those insights can help other re-searchers and practitioners, and inspire more interesting research in other fields, such as semantic segmentation, visual-linguistic pre-training, etc. One of the key findings is that both architectural changes and more epochs play critical roles in training vision transformer based detectors from scratch. Experiments on MS COCO datasets demonstrate that vision transformer based detectors trained from scratch can also achieve similar performances to their counterparts with ImageNet pre-training. Weixiang Hong 0001, Jiangwei Lao, Wang Ren, Jian Wang 0108, Jingdong Chen |
CVPR | 5 |
| 2022 | SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware NormalizationabstractRecently self-supervised representation learning has drawn considerable attention from the scene text recognition community. Different from previous studies using contrastive learning, we tackle the issue from an alternative perspective, i.e., by formulating the representation learning scheme in a generative manner. Typically, the neighboring image patches among one text line tend to have similar styles, including the strokes, textures, colors, etc. Motivated by this common sense, we augment one image patch and use its neighboring patch as guidance to recover itself. Specifically, we propose a Similarity-Aware Normalization (SimAN) module to identify the different patterns and align the corresponding styles from the guiding patch. In this way, the network gains representation capability for distinguishing complex patterns such as messy strokes and cluttered backgrounds. Experiments show that the proposed SimAN significantly improves the representation quality and achieves promising performance. Moreover, we surprisingly find that our self-supervised generative network has impressive potential for data synthesis, text image editing, and font interpolation, which suggests that the proposed SimAN has a wide range of practical applications. Canjie Luo, Jingdong Chen |
CVPR | 3 |
| 2022 | Hierarchical Memory Learning for Fine-Grained Scene Graph Generation
Youming Deng, Yansheng Li 0001, Yongjun Zhang 0002, Xiang Xiang 0001, Jian Wang 0108, Jingdong Chen, Jiayi Ma 0001 |
ECCV (27) | 6 |
| 2022 | The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And ResultsabstractIn this paper we discuss the rational of the Multi-model Information based Speech Processing (MISP) Challenge, and provide a detailed description of the data recorded, the two evaluation tasks and the corresponding baselines, followed by a summary of submitted systems and evaluation results. The MISP Challenge aims at tack-ling speech processing tasks in different scenarios by introducing information about an additional modality (e.g., video, or text), which will hopefully lead to better environmental and speaker robustness in realistic applications. In the first MISP challenge, two bench-mark datasets recorded in a real-home TV room with two reproducible open-source baseline systems have been released to promote research in audio-visual wake word spotting (AVWWS) and audio-visual speech recognition (AVSR). To our knowledge, MISP is the first open evaluation challenge to tackle real-world issues of AVWWS and AVSR in the home TV scenario. Hang Chen 0001, Hengshun Zhou, Jun Du 0002, Chin-Hui Lee 0001, Jingdong Chen, Shinji Watanabe 0001, Sabato Marco Siniscalchi, Odette Scharenborg, Diyuan Liu, Jianqing Gao, Cong Liu 0006 |
ICASSP | 5 |
| 2022 | DNN Based Multiframe Single-Channel Noise Reduction FiltersabstractWhile multiframe noise reduction filters, e.g., the multiframe Wiener and minimum variance distortionless response (MVDR) ones, have demonstrated great potential to improve both the subband and full-band signal-to-noise ratios (SNRs) by exploiting explicitly the interframe speech correlation, the implementation of such filters requires the knowledge of the interframe correlation coefficients for every subband, which are challenging to estimate in practice. In this work, we present a deep neural network (DNN) based method to estimate the interframe correlation coefficients and the estimated coefficients are subsequently fed into multiframe filters to achieve noise reduction. Unlike existing DNN based methods, which outputs the enhanced speech directly, the presented method combines deep learning and traditional methods, which gives more flexibility to optimize or tune noise reduction performance. Experimental results are presented to justify the properties of the presented methods. Ningning Pan, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2022 | Study of the Null Directions on The Performance of Differential BeamformersabstractNull directions are important parameters for differential beamformers, which play an important role on the beamforming performance. In this paper, we investigate the performance of differential beamformers as a function of the null directions. We first derive the directivity factor (DF) as an explicit function of null and show that the DF decreases to 0 if any null approaches to the desired look direction. We then validate the theoretical analysis through simulations using the beampattern, DF and signal-to-interference gain as the performance measures. The results show that: 1) the performance of a differential beamformer degrades significantly if there is any null close to the desired look direction; 2) with a fixed null direction, increasing the order of the differential beamformer can help improve performance. Xuehan Wang, Israel Cohen, Jacob Benesty, Jingdong Chen |
ICASSP | 4 |
| 2022 | Robust Pressure Matching with ATF Perturbation Constraints for Sound Field ControlabstractSound field control systems deployed in room acoustic environments require knowing the acoustic channel impulse responses between the loudspeakers and matching microphones, which are challenging to estimate accurately due to perturbations caused by such factors as temperature changes and sensors’ position mismatches. To deal with this issue, a robust pressure matching algorithm is developed in this work where a perturbation term of the acoustic transfer function (ATF) is modeled as a Gaussian process, based on which an uncertainty constraint is applied to limit the impact of perturbation on pressure matching. This constrained problem is formulated as one of biconvex optimization, and a coordinate descent algorithm is adopted to estimate the optimal control filter. Simulations are performed and results show that the proposed method is able to achieve more accurate control as compared to the standard pressure matching algorithm in the presence of ATF perturbations. Junqing Zhang, Liming Shi, Mads Græsbøll Christensen, Wen Zhang 0002, Lijun Zhang 0004, Jingdong Chen |
ICASSP | 6 |
| 2022 | Audio-Visual Speech Recognition in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we present the updated Audio-Visual Speech Recognition (AVSR) corpus of MISP2021 challenge, a large-scale audio-visual Chinese conversational corpus consisting of 141h audio and video data collected by far/middle/near microphones and far/middle cameras in 34 real-home TV rooms. To our best knowledge, our corpus is the first distant multi-microphone conversational Chinese audio-visual corpus and the first large vocabulary continuous Chinese lip-reading dataset in the adverse home-tv scenario. Moreover, we make a deep analysis of the corpus and conduct a comprehensive ablation study of all audio and video data in the audio-only/video-only/audiovisual systems. Error analysis shows video modality supplement acoustic information degraded by noise to reduce deletion errors and provide discriminative information in overlapping speech to reduce substitution errors. Finally, we also design a set of experiments such as frontend, data augmentation and end-to-end models for providing the direction of potential future work. The corpus and the code are released to promote the research not only in speech area but also for the computer vision area and cross-disciplinary research. Hang Chen 0001, Jun Du 0002, Yusheng Dai, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen |
INTERSPEECH | 8 |
| 2022 | Audio-Visual Wake Word Spotting in MISP2021 Challenge: Dataset Release and Deep AnalysisabstractIn this paper, we describe and release publicly the audio-visual wake word spotting (WWS) database in the MISP2021 Challenge, which covers a range of scenarios of audio and video data collected by near-, mid-, and far-field microphone arrays, and cameras, to create a shared and publicly available database for WWS. The database and the code 2 are released, which will be a valuable addition to the community for promoting WWS research using multi-modality information in realistic and complex conditions. Moreover, we investigated the different data augmentation methods for single modalities on an end-to-end WWS network. A set of audio-visual fusion experiments and analysis were conducted to observe the assistance from visual information to acoustic information based on different audio and video field configurations. The results showed that the fusion system generally improves over the single-modality (audio- or video-only) system, especially under complex noisy conditions. Hengshun Zhou, Jun Du 0002, Gongzhen Zou, Zhaoxu Nian, Chin-Hui Lee 0001, Sabato Marco Siniscalchi, Shinji Watanabe 0001, Odette Scharenborg, Jingdong Chen, Shifu Xiong, Jianqing Gao |
INTERSPEECH | 9 |
| 2022 | CRET: Cross-Modal Retrieval Transformer for Efficient Text-Video RetrievalabstractGiven a text query, the text-to-video retrieval task aims to find the relevant videos in the database. Recently, model-based (MDB) methods have demonstrated superior accuracy than embedding-based (EDB) methods due to their excellent capacity of modeling local video/text correspondences, especially when equipped with large-scale pre-training schemes like ClipBERT. Generally speaking, MDB methods take a text-video pair as input and harness deep models to predict the mutual similarity, while EDB methods first utilize modality-specific encoders to extract embeddings for text and video, then evaluate the distance based on the extracted embeddings. Notably, MDB methods cannot produce explicit representations for text and video, instead, they have to exhaustively pair the query with every database item to predict their mutual similarities in the inference stage, which results in significant inefficiency in practical applications. Kaixiang Ji, Jiajia Liu 0002, Weixiang Hong 0001, Liheng Zhong, Jian Wang 0108, Jingdong Chen |
SIGIR | 6 |
| 2022 | End-to-End Speaker Verification via Curriculum Bipartite Ranking Weighted Binary Cross-EntropyabstractEnd-to-end speaker verification achieves the verification through estimating directly the similarity score between a pair of utterances, which is formulated as a binary (i.e., target versus non-target) classification problem. Unlike the stage-wise method, an end-to-end verification approach optimizes the evaluation metrics directly and its output layer is parameter-free, which can save great computing and memory resources. However, there are two important issues that need to be meticulously handled in training an end-to-end speaker verification model. The first one is how to deal with severely imbalanced trials, i.e., the number of target trials is much smaller than that of nontarget trials, and the other is about how to handle easy trials that do not help improve the model in training. To circumvent these two issues, we propose in this paper a binary cross-entropy (BCE) type of loss function and present a method to train the deep neural network (DNN) models based on the proposed loss function for end-to-end speaker verification. The training process employs a bipartite ranking method to deal with the trial imbalance problem and a curriculum learning method to help improve both the training stability and performance of the model by selecting non-target trials from easy to hard ones gradually along the convergence process. Since the training process employs bipartite ranking and curriculum learning and the loss function is of the generalized BCE form, we name the new approach \textit{curriculum bipartite ranking weighted binary cross-entropy} (CBRW-BCE). Experimental results show that the model trained with CBRW-BCE not only achieves the state-of-the-art performance but is also well calibrated. Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Fundamental Approaches to Robust Differential Beamforming With High Directivity FactorsabstractDifferential beamforming, which measures the spatial derivatives of the acoustic pressure field, can be used in a wide range of small devices that require high-fidelity sound and speech acquisition as it can achieve frequency-invariant spatial responses with high directivity factors (DFs). Since a differential process is inherently sensitive to sensors' self noise and other array imperfections, the most challenging problem in the design of any differential beamformer is how to achieve the maximum possible DF while maintaining a proper level of robustness for practical usage. While significant efforts have been made on this topic, the problem remains unsolved and further study is indispensable. This paper is devoted to dealing with this challenging problem. It presents a study on theory and methods to achieve the optimal and fundamental compromise between the white noise gain (WNG), which quantifies how robust is the beamformer, and the DF in differential beamforming. The major contributions of this work are as follows. 1) We show and prove that any null constrained fixed beamformer can be decomposed as the sum of two orthogonal filters, i.e., the maximum WNG (MWNG) beamformer and a reduced-rank one. Based on this decomposition, we develop three kinds of differential beamformers from the WNG perspective, which can achieve a flexible and optimal compromise between DF and WNG. 2) We show that a transformed null constrained beamformer can also be decomposed as the sum of two orthogonal filters, i.e., the transformed maximum DF (MDF) beamformer and another reduced-rank one. Based on this decomposition, we also develop three kinds of differential beamformers, which can obtain the desired level of DF while using the rest of the degrees of freedom to maximize the WNG. Simulations are performed to validate the theoretical analysis and developed differential beamformers. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Kronecker Product Multichannel Linear Filtering for Adaptive Weighted Prediction Error-Based Speech DereverberationabstractReverberation, whichis caused by late reflections, impairs not only speech quality but also intelligibility. Consequently, dereverberation, a process to mitigate the impact of reverberation, has attracted significant research interests. Numerous approaches have been developed in the literature, among which the weighted-prediction-error (WPE) one has demonstrated promising potential for reducing or eliminating reverberation. The WPE method has been well studied and several variants have been developed. The adaptive one, called adaptive WPE (AWPE) method, has been widely investigated for use in real applications as it can deal with reverberation in time-varying acoustic environments. However, the computational complexity of AWPE is high, which may be a problem for its implementation in real-time systems. This paper presents some new insights into AWPE-based speech dereverberation by introducing the concepts of Kronecker product and partially time-varying filtering. It then develops two algorithms for dereverberation with lower complexity than AWPE. The significant contributions of this work are as follows. First, we propose a Kronecker product filtering framework for speech dereverberation, where the linear prediction filter is formulated as the Kronecker product of two sets of shorter filters. Second, we propose a partially time-varying Kronecker product filter for dereverberation. Instead of estimating the entire linear prediction filter as in the conventional method, the proposed one only needs to update part of the filter. The proposed approaches can significantly reduce the computational complexity without sacrificing dereverberation performance as compared to AWPE. Simulation results validate the theoretical analysis and justify the advantages of the new methods. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | On Differential Beamforming With Nonuniform Linear Microphone ArraysabstractWhile differential beamforming with uniform linear arrays (ULAs) has been widely studied, there is little work so far regarding the design of differential beamformers with nonuniform linear arrays (NULAs). This paper attempts to shed some light on the principles of differential beamforming with NULAs. We define spatial difference operators with NULAs, where any order of the spatial difference of the observation signals can be represented as the product of a nonuniform spatial difference operator matrix and the observation vector. Consequently, the design of differential beamformers is performed in two stages. In the first one, a nonuniform spatial difference operator matrix is applied to the array observations, thereby yielding differential signals. In the second stage, beamformers are designed and applied to the obtained differential signals to optimize the array performance. Based on the defined spatial difference operators, we derive from some performance metrics a family of differential beamformers with NULAs, which include the maximum directivity factor (DF), the maximum white noise gain (WNG), and the maximum front-to-back ratio (FBR) differential beamformers. To compromise between the DF and array robustness, we also derive the parameterized maximum DF and parameterized maximum FBR differential beamformers. The null-constraint maximum DF and WNG differential beamformers are also developed so that some nulls can be placed in specified directions for interference suppression. Simulation results validate the theoretical analysis and justify the properties of the proposed methods. Jilu Jin, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | A Framework of Directional-Gain Beamforming and a White-Noise-Gain-Controlled SolutionabstractIt is well known that an adaptive beamformer can be decomposed as a fixed beamformer followed by a post filter. This decomposition gives a much flexible way to design robust adaptive beamformers with high array gain and consequently it has become a popular approach to speech enhancement. In such a framework, the most critical problem is to estimate the post filter, which is studied in this paper. We present a multistage approach to the design of the post filter, which consists of a primary beamformer, a secondary beamformer and several auxiliary beamformers. Since the designed post filter is a function of source incidence angle by nature, we call it a directional gain. To evaluate the directivity of the beamformers in computing the directional gain, we introduce a modified beampattern, which is a function of both the source incidence angle and the point-source-to-background-noise ratio (PBR). To validate the presented approach, we analyze the principles of the primary, secondary and auxiliary beamformers, and then present a way to design these beamformers under the constraint of minimum white-noise-gain (WNG). Finally, we evaluate the performance of the directional gain and compare it with the traditional beamformers. The results show that the proposed directional gain can help achieve higher directivity factors (DFs), and better point-source-noise attenuation. Chao Pan 0001, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Microphone Array Beamforming With High Flexible Interference Attenuation and Noise ReductionabstractThis paper studies the problem of microphone array beamforming to enhance a speech signal of interest in adverse acoustic environments, where interference and additive background noise coexist. The problem is formulated as one of convex optimization whose solution under a specified level of interference attenuation leads to an interference controlled maximum noise reduction (ICMR) beamformer, which can be expressed as a linear combination of two MVDR beamformers: one attempts to extract the desired source signal while the other attempts to extract the interference. The combination coefficients are functions of the array manifold vectors, noise coherence matrix, and the specified interference attenuation factor. By tuning the interference attenuation factor, the ICMR beamformer can be implemented to achieve aggressive interference attenuation or even eliminate interference completely; but this may lead to less additive noise suppression or even noise amplification. To control the maximum sacrifice in gain (SG) of the signal-to-noise ratio (SNR) that is acceptable for additive reduction, a variant of ICMR is derived, which is named as the ICMR-SG beamformer. Simulations are performed and the results show that the ICMR beamformer is able to control the amount of interference attenuation. In comparison, ICMR-SG controls the maximum SG of SNR while achieving the optimal possible level of interference attenuation. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | CBNet: A Composite Backbone Network Architecture for Object Detectionabstracttop-performing object detectors depend heavily on backbone networks, whose advances bring consistent performance gains through exploring more effective network structures. In this paper, we propose a novel and flexible backbone framework, namely CBNet, to construct high-performance detectors using existing open-source pre-trained backbones under the pre-training fine-tuning paradigm. In particular, CBNet architecture groups multiple identical backbones, which are connected through composite connections. Specifically, it integrates the high- and low-level features of multiple identical backbone networks and gradually expands the receptive field to more effectively perform object detection. We also propose a better training strategy with auxiliary supervision for CBNet-based detectors. CBNet has strong generalization capabilities for different backbones and head designs of the detector architecture. Without additional pre-training of the composite backbone, CBNet can be adapted to various backbones (i.e., CNN-based vs. Transformer-based) and head designs of most mainstream detectors (i.e., one-stage vs. two-stage, anchor-based vs. anchor-free-based). Experiments provide strong evidence that, compared with simply increasing the depth and width of the network, CBNet introduces a more efficient, effective, and resource-friendly way to build high-performance backbone networks. Particularly, our CB-Swin-L achieves 59.4% box AP and 51.6% mask AP on COCO test-dev under the single-model and single-scale testing protocol, which are significantly better than the state-of-the-art results (i.e., 57.7% box AP and 50.2% mask AP) achieved by Swin-L, while reducing the training time by 6×. With multi-scale testing, we push the current best single model result to a new record of 60.1% box AP and 52.3% mask AP without using extra training data. Code is available at https://github.com/VDIGPKU/CBNetV2. Ting-Ting Liang, Xiaojie Chu, Yongtao Wang, Zhi Tang 0001, Jingdong Chen, Haibin Ling |
IEEE Trans. Image Process. | 7 |
| 2021 | LPSNet: A Lightweight Solution for Fast Panoptic SegmentationabstractPanoptic segmentation is a challenging task aiming to simultaneously segment objects (things) at instance level and background contents (stuff) at semantic level. Existing methods mostly utilize a two-stage detection network to attain instance segmentation results, and a fully convolutional network to produce a semantic segmentation prediction. Post-processing or additional modules are required to handle the conflicts between the outputs from these two nets, which makes such methods suffer from low efficiency, heavy memory consumption and complicated implementation. To simplify the pipeline and decrease computation/memory cost, we propose an one-stage approach called Lightweight Panoptic Segmentation Network (LPSNet), which does not involve a proposal, anchor or mask head. Instead, we predict a bounding box and semantic category at each pixel upon the feature map produced by an augmented feature pyramid, and design a parameter-free head to merge the per-pixel bounding box and semantic prediction into panoptic segmentation output. Our LPSNet is not only efficient in computation and memory, but also accurate in panoptic segmentation. Comprehensive experiments on COCO, Cityscapes and Mapillary Vistas datasets demonstrate the promising effectiveness and efficiency of the proposed LPSNet. Weixiang Hong 0001, Qingpei Guo, Jingdong Chen |
CVPR | 4 |
| 2021 | Planar Array Geometry Optimization for Region Sound AcquisitionabstractMicrophone arrays have been used in wide range of applications for sound acquisition and signal enhancement, the performance of which depends not only on the processing algorithms but also on the array geometry. A large number of efforts have been devoted to the development of beamforming and signal enhancement algorithms for processing microphone array signals in the literature. Relatively, few efforts have been made to investigate the problem of array geometry optimization. This paper studies the problem of geometry optimization for planar arrays and it develops a genetic optimization algorithm that can optimize the positions of the sensors, thereby maximizing the directivity factor (DF) with a constrained level of white noise gain (WNG) given the number of microphones, the region in which they should be placed, and the interested range of steering. Simulation results show that the optimized array geometry outperforms the uniform linear, the uniform circular and the rectangular grid geometries in terms of DF with the same number of sensors and the same constraint on the minimum level of WNG. Xi Chen 0128, Chao Pan 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2021 | Robust Recursive Least M-Estimate Adaptive Filter for the Identification of Low-Rank Acoustic SystemsabstractTo identify acoustic systems (which are low-rank in nature) in non-Gaussian and Gaussian noise, a robust recursive least M-estimate adaptive filtering algorithm is developed in this paper by applying the nearest Kronecker product to decompose the acoustic impulse response. Two M-estimators, i.e., the Cauchy and Welsch estimators, are employed to define the cost function of the adaptive filter, leading to a class of numerically stable adaptive filtering algorithms, which are robust to non-Gaussian noise. The effectiveness of the developed algorithm is validated in acoustic environments with both Gaussian and non-Gaussian noise. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 2 |
| 2021 | Combined Differential Beamforming With Uniform Linear Microphone ArraysabstractWhile differential beamformers have been widely used in voice communication and human-machine speech interface systems to enhance speech signals of interest, how to design such beamformers that on the one hand can achieve the highest possible directivity factor (DF) and on the other hand are able to obtain a certain level of white noise gain (WNG), so that they are robust enough to sensors’ self noise and array imperfections is still a challenging issue. This paper studies the problem of robust differential beamforming with small-size arrays to achieve a high DF. It presents a method for the design of differential beamformers with uniform linear arrays. We first generate differential pressure signals by applying the recently developed forward spatial difference operator to the outputs of the array with pressure sensors. The pressure microphone observation signals and the differential pressure signals are then put together, and a combined beamformer is subsequently designed, which consists of two subbeamformers, one operates on the pressure microphone observations and the other on the differential pressure signals. A new class of combined differential beamformers are introduced, which can achieve different levels of compromises between DF and WNG using an adjustable parameter. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
ICASSP | 5 |
| 2021 | Robust Steerable Differential Beamformers with Null Constraints for Concentric Circular Microphone ArraysabstractDifferential beamformers with concentric circular microphone arrays (CCMAs) are desirable for use in various applications since they can form frequency-invariant spatial responses, have better beam steering flexibility than linear arrays, and suffer less with beampattern irregularity and white noise amplification than circular microphone arrays (CMAs). The methods developed previously for differential beamforming with CCMAs are based on the series expansion. Such methods need to know the analytic form of the target beam-pattern, which may not be accessible in practice. Furthermore, expansion error may lead to erroneous solution, which can cause noise amplification instead of reduction. In this paper, we extend our recently developed beamforming method for CMAs to the design of differential beamformers with CCMAs, which takes advantage of the symmetric null constraints from the beampattern. Simulations are performed to justify the properties of the proposed approach. Xuehan Wang, Gongping Huang, Israel Cohen, Jacob Benesty, Jingdong Chen |
ICASSP | 5 |
| 2021 | A Simplified Wiener Beamformer Based on Covariance Matrix ModellingabstractThis paper is devoted to the problem of adaptive beamforming with small-spaced microphone arrays. In this context, the Wiener filter is an optimal beamformer in the mean-squared error (MSE) sense. However, it requires good estimates of the covariance matrices of the speech signal of interest and noise, which are difficult to achieve in time-varying and reverberant acoustic environments. To deal with this problem, we propose a general method by parametric modeling the covariance matrices of speech and noise, which leads to a simplified Wiener beamformer. This beamformer has only one time-varying parameter to estimate, which is much easier to achieve as compared to the estimation of covariance matrices. As an example, we adopt the parametric model used in the superdirective beamformer, which models the covariance matrices as a combination of the pseudo-coherence matrices of a point source and diffuse noise. Simulation results show that the developed beamformer outperforms the traditional Wiener beamformer in terms of both noise and reverberation suppression. Fan Zhang 0001, Chao Pan 0001, Jacob Benesty, Jingdong Chen |
ICASSP | 4 |
| 2021 | On the Design of Square Differential Microphone Arrays with a Multistage StructureabstractThis paper studies the problem of designing square differential microphone arrays (SDMAs). It presents a multistage approach, which first divides an SDMA composed of M2microphones into (M − 1)2subarrays with each subarray being a 2 × 2 square array formed by four adjacent microphones. Then, differential beamforming is performed with each subarray in the first-stage. The first-stage differential beamformers’ outputs are subsequently used as the inputs of the second stage to form (M − 2)2subarrays and a second-stage differential beamforming is then performed. Continuing this process till the (M −1)th stage, we obtain the final output of the SDMA. The SDMA designed in such a multistage structure has two important properties. First, the global weighting matrix is equal to the two dimensional convolution of weighting matrices from the first stage to the last one. Second, the global beampattern is equal to the product of beampatterns from all stages. Consequently, we can combine different kinds of beamformers in different stages and have better control of the performance metrics. Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
ICASSP | 4 |
| 2021 | Representation Learning of Remote Sensing Knowledge Graph for Zero-Shot Remote Sensing Image Scene ClassificationabstractAlthough deep learning has revolutionized remote sensing image scene classification, current deep learning-based approaches highly depend on the massive supervision of the predetermined scene categories and have disappointingly poor performance on new categories which go beyond the predetermined scene categories. In reality, the classification task often has to be extended along with the emergence of new applications that inevitably involve new categories of remote sensing image scenes, so how to make the deep learning model own the inference ability to recognize the remote sensing image scenes from unseen categories becomes incredibly important. By fully exploiting the remote sensing domain characteristic, this paper proposes a novel remote sensing knowledge graph-guided deep alignment network to address zero-shot remote sensing image scene classification. To improve the semantic representation ability of remote sensing-oriented scene categories, this paper, for the first time, tries to generate the semantic representations of remote sensing scene categories by representation learning of remote sensing knowledge graph (SR-RSKG). In addition, this paper proposes a novel deep alignment network with a series of constraints (DAN) to conduct robust cross-modal alignment between visual features and semantic representations. Extensive experiments on one merged remote sensing image scene dataset, which is the integration of multiple publicly open remote sensing image scene datasets, show that the presented SR-RSKG obviously outperforms the existing semantic representation methods (e.g., the natural language processing models and manually annotated attribute vectors), and our proposed DAN shows better performance compared with the state-of- the-art methods under different kinds of semantic representations. Yansheng Li 0001, Yongjun Zhang 0002, Ruixian Chen, Jingdong Chen |
IGARSS | 5 |
| 2021 | MatchVIE: Exploiting Match Relevancy between Entities for Visual Information ExtractionabstractVisual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify each kind of semantics by introducing multimodal features, such as font, color, layout. But simply introducing multimodal features can't work well when faced with numeric semantic categories or some ambiguous texts. To address this issue, in this paper we propose a novel key-value matching model based on a graph neural network for VIE (MatchVIE). Through key-value matching based on relevancy evaluation, the proposed MatchVIE can bypass the recognitions to various semantics, and simply focuses on the strong relevancy between entities. Besides, we introduce a simple but effective operation, Num2Vec, to tackle the instability of encoded values, which helps model converge more smoothly. Comprehensive experiments demonstrate that the proposed MatchVIE can significantly outperform previous methods. Notably, to the best of our knowledge, MatchVIE may be the first attempt to tackle the VIE task by modeling the relevancy between keys and values and it is a good complement to the existing methods. Guozhi Tang, Lele Xie, Jingdong Chen, Qianying Wang 0002, Yaqiang Wu |
IJCAI | 5 |
| 2021 | AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference ScenarioabstractIn this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario. The dataset consists of 211 recorded meeting sessions, each containing 4 to 8 speakers, with a total length of 120 hours. This dataset aims to bridge the advanced research on multi-speaker processing and the practical application scenario in three aspects. With real recorded meetings, AISHELL-4 provides realistic acoustics and rich natural speech characteristics in conversation such as short pause, speech overlap, quick speaker turn, noise, etc. Meanwhile, accurate transcription and speaker voice activity are provided for each meeting in AISHELL-4. This allows the researchers to explore different aspects in meeting processing, ranging from individual tasks such as speech front-end processing, speech recognition and speaker diarization, to multi-modality modeling and joint optimization of relevant tasks. Given most open source dataset for multi-speaker tasks are in English, AISHELL-4 is the only Mandarin dataset for conversation speech, providing additional value for data diversity in speech community. We also release a PyTorch-based training and evaluation framework as baseline system to promote reproducible research in this field. Yihui Fu, Luyao Cheng, Shubo Lv, Yukai Jv, Yuxiang Kong, Zhuo Chen 0006, Yanxin Hu, Lei Xie 0001, Jian Wu 0027, Hui Bu, Jun Du 0002, Jingdong Chen |
Interspeech | 13 |
| 2021 | A Conceptual Approach of Passive Human-Intention-Orientated Variable Admittance Control using Power EnvelopeabstractTwo main challenges that need to be addressed in physical human-robot interaction (pHRI) are efficient recognition of human intention and interaction safety. In this paper, a general human intention framework was summarized, firstly, according to the robot's roles: a passive follower and a compliant leader. Secondly, we proposed variable admittance control models governed by human intentions. Power envelope approaches were then proposed to impose constraints on the variable admittance parameters inferred from human intention for maintaining passivity conservatively. Our passivity preserving approaches were validated via simulation and shown to avoid mismatching of time-varying admittance parameters that restrain drastic changes of admittance controller dynamics, which usually result in instability. Finally, the relationship between the robot's passivity and stability when it interacts with the human was analyzed. Jingdong Chen, Paul I. Ro |
IROS | 1 |
| 2021 | GilBERT: Generative Vision-Language Pre-Training for Image-Text RetrievalabstractGiven a text/image query, image-text retrieval aims to find the relevant items in the database. Recently, visual-linguistic pre-training (VLP) methods have demonstrated promising accuracy on image-text retrieval and other visual-linguistic tasks. These VLP methods are typically pre-trained on a large amount of image-text pairs, then fine-tuned on various downstream tasks. Nevertheless, due to the natural modality incompleteness in image-text retrieval, i.e., the query is either image or text rather than an image-text pair, the naive application of VLP to image-text retrieval results in significant inefficiency. Moreover, existing VLP methods cannot extract comparable representations for a single-modal query and multi-modal database items. In this work, we propose a generative visual-linguistic pre-training approach, termed as GilBERT, to simultaneously learn generic representations of image-text data and complete the missing modality for incomplete pairs. In testing phase, the proposed GilBERT facilitates efficient vector-based retrieval by providing unified feature embedding for query and database items. Moreover, the generative training not only makes GilBERT compatible with non-parallel text/image corpus, but also enables GilBERT to model the image-text relationships without suffering massive randomly-sampled negative samples, leading to superior experimental performances. Extensive experiments demonstrate the advantages of GilBERT in image-text retrieval, in terms of both efficiency and accuracy. Weixiang Hong 0001, Kaixiang Ji, Jiajia Liu 0002, Jian Wang 0108, Jingdong Chen |
SIGIR | 5 |
| 2021 | A New Method to Design Steerable First-Order Differential BeamformersabstractFirst-order differential microphone arrays (FODMAs), which combine a small-spacing uniform linear array and a first-order differential beamformer, have been used in a wide range of applications for sound and speech signal acquisition. However, traditional FODMAs are not steerable and their main lobe can only be at the endfire directions. To circumvent this problem, we propose in this letter a new method to design steerable FODMAs. We first divide the target beampattern into a sum of two sub-beampatterns, i.e., cardioid and dipole, where the summation is controlled by the steering angle. We then design two sub-beamformers, one is similar to the traditional approach and is used to achieve the cardioid sub-beampattern, while the other is designed to filter the squared observation signals and is used to approximate the dipole sub-beampattern. The overall beampattern resembles the target beampattern for any steering angle. Simulations and experiments are performed to justify the effectiveness of the developed method. Xin Leng, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Single-Input/Binaural-Output Antiphasic Speech Enhancement Method for Speech Intelligibility ImprovementabstractImproving intelligibility of a speech signal of interest from its observations (with a single microphone) corrupted by additive noise has long been a challenging problem. Motivated by important findings achieved in the psychoacoustic field, we propose in this work a deep learning based method to render the noise and desired speech in the perceptual space such that the perception of the desired speech is least affected by the noise. Specifically, we adopt the temporal convolutional network (TCN) based structure to map the single-channel noisy observations into two binaural signals, one for the left ear and the other for the right ear. The TCN is trained in such a way that the desired speech and noise will be perceived to be in opposite directions when the listener listens to the binaural signals. This antiphasic binaural presentation enables the listener to better distinguish the desired speech from the annoying noise for improved speech intelligibility. The modified rhyme test is performed for evaluation and the results justify the superiority of the proposed method for speech intelligibility improvement. Ningning Pan, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 3 |
| 2021 | Time Difference of Arrival Estimation Based on a Kronecker Product DecompositionabstractTime difference of arrival (TDOA) estimation, which often serves as the fundamental step for a source localization or a beamforming system, has a significant practical importance in a wide spectrum of applications. To deal with reverberation, the TDOA estimation problem is often transformed into one of identifying the relative acoustic impulse responses. This letter presents a method to efficiently identify the relative acoustic impulse response between two microphones for TDOA estimation based on the so-called Kronecker product decomposition. By decomposing the relative impulse response into a series of Kronecker products of shorter filters, the original channel identification problem with a long impulse response is converted into one of identifying a number of short filters. Since the TDOA information is embedded only in the direct path of the relative impulse response, the dimension of the Kronecker product decomposition can be very small and, as a result, the developed algorithm is expected to work well in real environments with a small number of data snapshots. Xianrui Wang, Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
IEEE Signal Process. Lett. | 4 |
| 2021 | Robust Dereverberation With Kronecker Product Based Multichannel Linear PredictionabstractReverberation impairs not only the speech quality, but also intelligibility. The weighted-prediction-error (WPE) method, which estimates the late reverberation component based on a multichannel linear predictor, is by far one of the most effective algorithms for dereverberation. Generally, the WPE prediction filter in every short-time-Fourier-transform (STFT) subband has to be long enough to estimate accurately the late reverberation component. As a consequence, WPE is computationally expensive, which makes it difficult to implement into real-time embedded or edge computing devices. Moreover, WPE is sensitive to additive noise and its performance may suffer from dramatic degradation even in environments where the signal-to-noise ratio (SNR) is high. To address these drawbacks, this letter proposes to decompose the multichannel linear prediction filter as a Kronecker product of a temporal (interframe) prediction filter and a spatial filter. An iterative algorithm is then developed to optimize the two filters. In comparison with the original WPE algorithm, the presented method not only exhibits better performance in terms of dereverberation and robustness to additive noise, as there are fewer parameters to estimate for a given number of observation signal samples, but is also computationally more efficient, since the dimensions of the covariance matrices after Kronecker product decomposition are smaller. Wenxing Yang, Gongping Huang, Jingdong Chen, Jacob Benesty, Israel Cohen, Walter Kellermann |
IEEE Signal Process. Lett. | 3 |
| 2021 | On a Particular Family of Differential Beamformers With Cardioid-Like and No-Null PatternsabstractDifferential microphone arrays (DMAs), which are responsive to the differential acoustic pressure fields, have been used in a wide range of applications related to audio and speech. The core part of a DMA is the so-called differential beamformer, which is generally designed by placing a number of nulls in its beampattern to attenuate noise from some directions. But the presence of these nulls may cause some great issues, e.g., leading to suboptimal performance if the interference/noise is incident from directions other than the nulls' directions, and making the beamformer less robust to sensors' self noise and array imperfections. To overcome these problems, this letter is devoted to the design of differential beamformers with no nulls in its beampattern. A design method and its multistage implementation are presented and analyzed. An improved solution is then developed, which is able to form frequency-invariant beampatterns with no nulls in the frequency range of speech signals. Simulations are provided to illustrate the properties of the developed methods. Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE Signal Process. Lett. | 4 |
| 2021 | On the Robustness of the Superdirective BeamformerabstractIn microphone array beamforming, a high directional gain is always desired for acoustic noise and reverberation suppression; as a result, the superdirective beamformer has been of great interest in many applications. However, this beamformer is well known to be very sensitive to array imperfections. While much effort has been made to improve its robustness, it is still a major problem. This paper is essentially devoted to the study of the robustness of the superdirective beamformer and derivation of better ways to deal with this important issue. We first prove that any distortionless fixed beamformer can be written as the sum of two orthogonal beamformers, i.e., the sum of the classical delay-and-sum (DS) beamformer and a reduced-rank beamformer. Based on this property, different kinds of robust superdirective beamformers are then developed. We also show that the robust design problem can be transformed into a quadratic eigenvalue problem (QEP), which leads to a solution that achieves the maximum possible directivity factor (DF) while meets the white noise gain (WNG) constraint over a frequency band of interest. Xi Chen 0128, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Steering Study of Linear Differential Microphone ArraysabstractDifferential microphone arrays (DMAs) can achieve high directivity and frequency-invariant spatial response with small apertures; they also have a great potential to be used in a wide spectrum of applications for high-fidelity sound acquisition. Although many efforts have been made to address the design of linear DMAs (LDMAs), most developed methods so far only work for the situation where the source of interest is incident from the endfire direction. This paper studies the steering problem of differential beamformers with linear microphone arrays. We present new insights into beam steering of LDMAs and propose a series of steerable differential beamformers. The major contributions of this paper are as follows. 1) A series of ideal functions are defined to describe the ideal, target beampatterns of LDMAs. 2) We prove that first-order differential beamformers with linear microphone arrays are not steerable and their mainlobes can only be at the endfire directions. 3) We deduce the fundamental conditions for designing steerable differential beamformers with LDMAs. 4) We develop a method to design steerable beamformers with LDMAs using null constraints. Simulations and experiments validate the properties of the developed method. Jilu Jin, Gongping Huang, Xuehan Wang, Jingdong Chen, Jacob Benesty, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Beamforming with Cube Microphone Arrays Via Kronecker Product DecompositionsabstractMicrophone arrays combined with beamforming have been widely used to solve many important acoustic problems in a wide range of applications. Much effort has been devoted in the literature to microphone array beamforming, among which the Kronecker product beamforming method developed recently has demonstrated some interesting properties. Generally, this method decomposes the global beamforming filter into a Kronecker product of a number of sub-beamforming filters, each of which corresponds to a virtual subarray and can be designed individually. This decomposition not only reduces significantly the number of beamforming coefficients, but also can be explored to improve the robustness and flexibility of beamforming. This paper extends Kronecker product beamforming from two-dimensional arrays into three-dimensional cube arrays. We consider two decompositions, i.e., fully and partially separable ones. The former decomposes the entire array into three linear subarrays while the latter decomposes the entire array into a linear subarray and a planar one. Then, for each case, we derive the Kronecker product maximum white noise gain beamformer, the Kronecker product approximate maximum directivity factor (DF) beamformer, the Kronecker product null-steering beamformer, and the Kronecker product iterative maximum DF beamformer. Simulation results demonstrate the properties and advantages of the proposed beamformers. Xuehan Wang, Jacob Benesty, Jingdong Chen, Gongping Huang, Israel Cohen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | A New Class of Differential BeamformersabstractDifferential microphone arrays (DMAs) have been used in a wide range of applications for high-fidelity acoustic signal acquisition and enhancement. In the design of differential beamformers, three of the widely used measures are the directivity factor (DF), the front-to-back ratio (FBR), and the white noise gain (WNG). The former two have been used to obtain optimal differential beamformers, e.g., the hypercardioid and supercardioid, and the third one is generally used to analyze and control the robustness of the beamformer with respect to array imperfections due to sensors' self noise, mismatch among sensors, and sensors' placement errors. In this paper, we present a new measure called directivity factor and front-to-back ratio (DFBR), which is a generalization of DF and FBR. With this new measure, three different kinds of beamformers are derived. The first one is the maximum DFBR beamformer, which is deduced by maximizing DFBR with a joint diagonalization method. The second one is the ψ-cardioid beamformer, which is the maximum DFBR beamformer corresponding to a distortionless constraint. The last one is the reduced-rank differential beamformer, which is obtained by properly choosing the dimension of the signal subspace and maximizing WNG subject to the distortionless constraint. The developed beamformers have many interesting properties, which are justified by both simulations and experiments. Wenxing Yang, Jacob Benesty, Gongping Huang, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Differential Beamforming From the Beampattern Factorization PerspectiveabstractDifferential beamformers have demonstrated a great potential in forming frequency-invariant beampatterns and achieving high directivity factors. Most conventional approaches design differential beamformers in such a way that their beampatterns resemble a desired or target beampattern. In this paper, we show how to design differential beamformers by simply taking advantage of the fact that the beampattern is actually a particular form of an exponential polynomial. Thanks to this quite obvious formulation, a target beampattern is not really needed while the zeros of the exponential polynomial and/or its factorization are fully exploited. The advantage of this factorization is twofold. First, it gives the relation between the beamformer and the roots of the polynomial, so the former can be directly determined from the latter, which seems natural and convenient. Second, based on this factorization, we propose a new formulation of the beamforming filter, which decomposes the filter into shorter ones with the Kronecker product. This formulation is very general, and many well-known beamformers such as the differential and delay-and-sum (DS) ones can be derived from it. Furthermore, the new formulation allows one to combine different kinds of beamformers together, which gives a great flexibility in forming different beampatterns and achieving a compromise among the directivity factor (DF), white noise gain (WNG), and frequency invariance. Jacob Benesty, Jingdong Chen, Gongping Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | On the Design of 3D Steerable Beamformers With Uniform Concentric Circular Microphone ArraysabstractCircular microphone arrays (CMAs) and concentric CMAs (CCMAs) have been used in a wide range of applications such as smartspeakers and teleconferencing systems because of their flexible steering ability. Although many efforts have been devoted to beamforming with CCMAs, most existing methods consider only the 2-dimensional (2D) case and assume that the sound sources of interest are in the same plane as the sensor array (generally the horizontal plane), which often does not hold true in practical applications. This paper deals with the problem of beamforming with uniform CCMAs (UCCMAs) in the 3-dimensional (3D) space to control the steering of the spatial response and meanwhile form frequency-invariant beampatterns for processing broadband acoustic and speech signals. The major contributions of this work are summarized as follows: 1) it presents an analysis based on the spherical harmonics decomposition about the Nth-order optimal and steerable directivity patterns; 2) a beamforming method is developed in which the beamformer's coefficients are identified by solving a linear system of equations formed by approximating the Nth-order optimal target beampattern with the beamformer's beampattern while the resulting beampattern can be steered flexibly in the 3D space; 3) the sufficient and necessary condition on the array geometry and sensors' placement are given to ensure that the beamformer exists and is unique; and 4) the analytical forms of the directivity factor (DF) and white noise gain (WNG) of the resulting beamformer is given and discussion is presented on what conditions irregularities (deep nulls) in WNG and DF may occur. Simulations are provided to illustrate the property of the developed beamforming methods. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Variational Connectionist Temporal Classification
Linlin Chao, Jingdong Chen |
ECCV (28) | 2 |
| 2020 | Partial AUC Optimization Based Deep Speaker Embeddings with Class-Center Learning for Text-Independent Speaker VerificationabstractDeep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and identification. The verification loss functions match the pipeline of speaker verification, but their implementations are difficult. Thus, most state-of-the-art deep embedding methods use the identification loss functions with softmax output units or their variants. In this paper, we propose a verification loss function, named the maximization of partial area under the Receiver-operating-characteristic (ROC) curve (pAUC), for deep embedding based text-independent speaker verification. We also propose a class-center based training trial construction method to improve the training efficiency, which is critical for the proposed loss function to be comparable to the identification loss in performance. Experiments on the Speaker in the Wild (SITW) and NIST SRE 2016 datasets show that the proposed pAUC loss function is highly competitive with the state-of-the-art identification loss functions. Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen |
ICASSP | 3 |
| 2020 | Robust Frequency-Domain Recursive Least M-Estimate Adaptive Filter For Acoustic System IdentificationabstractTo identify acoustic systems in non-Gaussian and Gaussian noises, a robust frequency-domain recursive least M-estimate (FRLM) adaptive filtering algorithm is proposed. The cost function of the adaptive filter is defined by using a robust time-domain M-estimator, while its update equation is derived from the normal equation in the frequency domain. As compared to the frequency-domain recursive least-squares adaptive filter, the FRLM algorithm obtains the robustness to non-Gaussian and Gaussian noises. The performance of the proposed algorithm is validated in simulated acoustic environments. Hongsen He, Jingdong Chen, Jacob Benesty, Yi Yu 0002 |
ICASSP | 2 |
| 2020 | Robust and steerable kronecker product differential beamforming With rectangular microphone arraysabstractDifferential microphone arrays (DMAs), a class of welldesigned small-size arrays combined with differential beamforming, are very useful for processing broadband acoustic, audio, and speech signals in a wide range of applications. However, most efforts in the literature so far have been devoted to linear, circular, and spherical arrays. In this paper, we consider rectangular shapes of planar microphone arrays. Instead of adopting the traditional differential beamforming methods developed in the literature, we present a differential beamforming method based on the so-called Kronecker product. We first decompose the entire rectangular array into two virtual rectangular sub-arrays so that the steering vector of the entire array is the Kronecker product of the steering vectors of the two smaller virtual rectangular sub-arrays. We use the first virtual rectangular array, which is much smaller in size than the entire array but well satisfies the basic requirements for differential beamforming, to design a steerable differential beamformer. For the second virtual rectangular array, we can design either the delay-and-sum (DS) beamformer, which helps to improve the robustness of the global differential beamformer, or an adaptive beamformer, which makes the global differential beamformer adaptive. This method has many interesting properties, particularly the designed beamformer is fully steerable, and its robustness and the array gain can be easily controlled. Gongping Huang, Jacob Benesty, Jingdong Chen, Israel Cohen |
ICASSP | 3 |
| 2020 | Proximal Multitask Learning Over Distributed Networks with Jointly Sparse StructureabstractModeling relations between local optimum parameter vectors in multitask networks has attracted much attention over the last years. This work considers a distributed optimization problem for parameter vectors with a jointly sparse structure among nodes, that is, the parameter vectors share the same support set. By introducing an ℓ∞,1-norm penalty at each node, and using a proximal gradient method to minimize the regularized cost, we devise a proximal multitask diffusion LMS algorithm which promotes the joint-sparsity to enhance the estimation performance. Analyses are provided to ensure the stability. Simulation results are presented to highlight the performance. Danqi Jin, Jie Chen 0022, Cédric Richard, Jingdong Chen |
ICASSP | 4 |
| 2020 | An Improved Solution to the Frequency-Invariant Beamforming with Concentric Circular Microphone ArraysabstractFrequency-invariant beamforming with circular microphone arrays (CMAs) has drawn a significant amount of attention for its steering flexibility and high directivity. However, frequency-invariant beam-forming with CMAs often suffers from the so-called null problem, which is caused by the zeros of the Bessel functions; then, concentric CMAs (CCMAs) are used to deal with this problem. While frequency-invariant beamforming with CCMAs can mitigate the null problem, the beampattern is still suffering from distortion due to s-patial aliasing at high frequencies. In this paper, we find that the spatial aliasing problem is caused by higher-order circular harmonics. To deal with this problem, we take the aliasing harmonics into account and approximate the beampattern with a higher truncation order of the Jacobi-Anger expansion than required. Then, the beam-forming filter is determined by minimizing the errors between the desired directivity pattern and the approximated one. Simulation results show that the developed method can mitigate the distortion of the beampattern caused by spatial aliasing. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2020 | A class of multichannel sparse linear prediction algorithms for time delay estimation of speech sources
Hongsen He, Jingdong Chen, Jacob Benesty, Wenxing Zhang, Tao Yang 0039 |
Signal Process. | 2 |
| 2020 | Generalized combined nonlinear adaptive filters: From the perspective of diffusion adaptation over networks
Wenxia Lu, Lijun Zhang 0004, Jie Chen 0022, Jingdong Chen |
Signal Process. | 4 |
| 2020 | Cosine metric learning based speaker verification
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen |
Speech Commun. | 3 |
| 2020 | Beamforming With Small-Spacing Microphone Arrays Using Constrained/Generalized LASSOabstractIn this letter, we develop an approach to the design of beamformers with small-spacing uniform linear microphone arrays by incorporating sparseness constraints for attenuating scattered interference incident from some pre-specified ranges of directions of arrival. The design process is formulated as a constrained LASSO problem. By adjusting the value of a tuning parameter, the proposed method can make compromises among three important yet conflicting (especially at low frequencies) performance measures of small-spacing microphone arrays, i.e., the directivity factor (DF), which quantifies the array spatial gain, the white noise gain (WNG), which evaluates the robustness of the beamformer, and the signal-to-interference-ratio (SIR) gain with respect to scattered interference. Simulation results illustrate the properties of the developed approach. Xianghui Wang, Jacob Benesty, Jingdong Chen, Israel Cohen |
IEEE Signal Process. Lett. | 3 |
| 2020 | Speaker Verification by Partial AUC Optimization With Mahalanobis Distance Metric LearningabstractReceiver operating characteristic (ROC) and detection error tradeoff (DET) curves are two widely used evaluation metrics for speaker verification. They are equivalent since the latter can be obtained by transforming the former's true positive y-axis to false negative y-axis and then re-scaling both axes by a probit operator. Real-world speaker verification systems, however, usually work on part of the ROC curve instead of the entire ROC curve given an application. Therefore, we propose in this article to use the area under part of the ROC curve (pAUC) as a more efficient evaluation metric for speaker verification. A Mahalanobis distance metric learning based back-end is applied to optimize pAUC, where the Mahalanobis distance metric learning guarantees that the optimization objective of the back-end is a convex one so that the global optimum solution is achievable. To improve the performance of the state-of-the-art speaker verification systems by the proposed back-end, we further propose two feature preprocessing techniques based on length-normalization and probabilistic linear discriminant analysis respectively. We evaluate the proposed systems on the major languages of NIST SRE16 and the core tasks of SITW. Experimental results show that the proposed back-end outperforms the state-of-the-art speaker verification back-ends in terms of seven evaluation metrics. Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Differential Beamforming on GraphsabstractWe study differential beamforming from a graph perspective. The microphone array used for differential beamforming is viewed as a graph, where its sensors correspond to the nodes, the number of microphones corresponds to the order of the graph, and linear spatial difference equations among microphones are related to graph edges. Specifically, for the first-order differential beamforming with an array of M microphones, each pair of adjacent microphones are directly connected, resulting in M - 1 spatial difference equations. On a graph, each of these equations corresponds to a 2-clique. For the second-order differential beamforming, each three adjacent microphones are directly connected, resulting in M - 2 second-order spatial difference equations, and each of these equations corresponds to a 3-clique. In an analogous manner, the differential microphone array for any order-of-differential beamforming can be viewed as a graph. From this perspective, we then derive a class of differential beamformers, including the maximum white noise gain beamformer, the maximum directivity factor one, and optimal compromising beamformers. Simulations are presented to demonstrate the performance of the derived differential beamformers. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | A Simple Theory and New Method of Differential Beamforming With Uniform Linear Microphone ArraysabstractThis article presents a theoretical study of differential beamforming with uniform linear arrays. By defining a forward spatial difference operator, any order of the spatial difference of the observed signals can be represented as a product of a difference operator matrix and the microphone array observations. Consequently, differential beamforming is implemented in two stages, where the first one obtains spatial difference of the observations and the second stage optimizes the beamformer. The major contributions of this article are as follows. First, we propose a new theory of differential beamforming with uniform linear arrays, which shows clearly the connection between the conventional differential beamforming and the null-constrained differential beamforming methods. This provides some new insight into the design of differential beamformers. Second, we deduce some new differential beamformers, where conventional beamforming may be seen as a particular case. Specifically, we derive the maximum white noise gain (MWNG), maximum directivity factor (MDF), parameterized MDF, and parameterized maximum front-to-back ratio differential beamformers. Third, we further extend the idea of how to design optimal differential beamformers by combining both the observed signals and their spatial differences. Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Design of Planar Differential Microphone Arrays With Fractional OrdersabstractDifferential microphone arrays (DMAs) often encounter white noise amplification, especially at low frequencies. If the array geometry and the number of microphones are fixed, one can improve the white noise amplification problem by reducing the DMA order. With the existing differential beamforming methods, the DMA order can only be a positive integer number. Consequently, with a specified beampattern (or a kind of beampattern), reducing this order may easily lead to over compensation of the white noise gain (WNG) and too much reduction of the directivity factor (DF), which is not optimal. To deal with this problem, we present in this article a general approach to the design of DMAs with fractional orders. The major contributions of this article include but are not limited to: 1) we first define a directivity pattern that can achieve a continuous compromise between the pattern corresponding to the maximum DMA order and the omnidirectional pattern; 2) by approximating the beamformer's beampattern with the Jacobi-Anger expansion, we present a method to find the proper differential beamforming filter so that its beampattern matches closely the target directivity pattern of fractional orders; and 3) we show how to determine analytically the proper fractional order of the DMA with a given target beampattern when either the value of the DF or WNG is specified, which is useful in practice to achieve the desired beampattern and spatial gain while maintaining the robustness of the DMA system. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | On Estimation of Time-Varying Variances of Source and Noise for Sensor Array ProcessingabstractEstimation of time-varying variances of signals for beamforming in sensor arrays is a challenging problem. Based on the assumption that the array manifold vector and the noise pseudo-coherence matrix are known a priori or are well estimated, we present in this paper two estimators for estimating the time-varying variances of the source signal of interest and the noise. These two estimators are then extended to deal with the following situations: 1) there are multiple candidates of the noise pseudo-coherence matrix or the noise pseudo-coherence matrix is a linear combination of some base pseudo-coherence matrices, and 2) the estimation variance is large and smoothing is needed. Simulations for speech enhancement applications are performed and the results show that the proposed estimators can well track the time-varying variances of both the speech and noise signals. It is also demonstrated that the optimal beamformer using the variance parameters estimated with the presented estimators outperforms the widely used traditional optimal beamformers in terms of improvement in both the signal-to-noise ratio (SNR) and the log-spectral distortion (LSD). Chao Pan 0001, Jingdong Chen, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | AUC Optimization for Deep Learning Based Voice Activity DetectionabstractVoice activity detection (VAD) based on deep neural networks (DNN) has demonstrated good performance in adverse acoustic environments. Current DNN based VAD optimizes a surrogate function, e.g. minimum cross-entropy or minimum squared error, at a given decision threshold. However, VAD usually works on-the-fly with a dynamic decision threshold; and ROC curve is a global evaluation metric of VAD that reflects the performance of VAD at all possible decision thresholds. In this paper, we propose to optimize the area under ROC curve (AUC) by DNN, which can maximize the performance of VAD in terms of the ROC curve. Experimental results show that optimizing AUC by DNN results in higher performance than the common method of optimizing the minimum squared error by DNN. Zi-Chen Fan, Zhongxin Bai, Xiao-Lei Zhang 0001, Susanto Rahardja, Jingdong Chen |
ICASSP | 5 |
| 2019 | Properties and Limits of the Minimum-norm Differential Beamformers with Circular Microphone ArraysabstractSmall aperture circular microphone arrays (CMAs) have been widely used in many applications such as teleconferencing, smartspeakers, and robotics. A critical component of such arrays is the differential beamformer, which can achieve relatively high spatial gains with the same beampatterns at most frequencies. Among different differential beamforming approaches that were developed in the literature, the minimum-norm one has attracted much interest as it can deal better with sensors' self noise, sensor mismatch, and beamformer's irregularity at some frequencies due to the zeros of the Bessel functions. In our previous study, we have investigated the performance of the minimum-norm differential beamformer with uniform CMAs (UCMAs) in the 2-dimensional (2D) space where the sound sources and the sensors are assumed to be in the same plane. But in practice, this assumption is generally not true. So, in this paper, we investigate the properties and limitations of the minimum-norm differential beamformer in the 3-dimensional(3D) space. Through theoretical study as well as simulations, we show that the minimum-norm differential beamformer is effective in dealing with the problem of white noise amplification and irregularity of the beampatterns and the directivity factor (DF) if the steering angles are within or near the sensor plane, but it becomes less and less effective as the beamformer is steered away from this plane. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2019 | Design of Optimal Linear Differential Microphone Arrays Based Array Geometry OptimizationabstractThis paper presents a method to design optimal linear differential microphone arrays (DMAs) by optimizing the array geometry. By constraining the DMA beamformer to achieve a given target value of the directivity factor (DF) with a specified target frequency-invariant beampattern while achieving also the highest possible white noise gain (WNG), an optimization algorithm is developed, which consists of the following two steps. 1) The full frequency band of interest is divided into a few subbands. At every subband, the entire linear array is divided into subarrays and the number of subarrays depends on the total number of the sensors and the order of the DMA. A cost function is then defined, which is minimized to determine what subarray produces the optimal performance. 2) The subband optimal subarrays are then combined across the entire frequency band to form a fullband cost function, from which the geometry of the entire array is optimized. These two steps are repeated with the particle swarm optimization (PSO) algorithm until the desired array performance is reached. Simulation results demonstrate that the proposed method can obtain the target DF with a frequency-invariant beampattern over a wide band of frequencies while maintaining a reasonable level of WNG. Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 3 |
| 2019 | On the Design of Flexible Kronecker Product Beamformers with Linear Microphone ArraysabstractThis paper proposes a method for the design of flexible Kronecker product beamformers based on the decomposition of the steering vector of a physical array as a Kronecker product of steering vectors of two smaller virtual arrays. With this decomposition, the global beamforming filter is designed by optimizing the two sub-beamformers in a cascaded manner, which can offer much flexibility to control the performance of beamforming or control the compromise between different, conflicted performance measures. In comparison with a recently developed method that restricts the number of microphones of the given physical array to a multiplication of two integers, each corresponding to the number of sensors of one virtual array, the approach in this work decomposes the physical array in such a way that the sensors in the two virtual arrays may share positions and the number of microphones of the physical array can be any positive integer. Simulations demonstrate the properties of the proposed approach. Wenxing Yang, Gongping Huang, Jacob Benesty, Israel Cohen, Jingdong Chen |
ICASSP | 5 |
| 2019 | An iterative mask estimation approach to deep learning based multi-channel speech recognition
Yanhui Tu, Jun Du 0002, Lei Sun 0010, Hai-Kun Wang, Jingdong Chen, Chin-Hui Lee 0001 |
Speech Commun. | 6 |
| 2019 | Recursive Variable Span Linear Filter for Noise ReductionabstractThe design of variable span linear filters for noise reduction involves a generalized eigenvalue decomposition problem that is of high computational complexity. In order to address this issue, this work proposes a recursive algorithm that computes the filter weights with streaming signal data. Specifically, the inverse square root of the noise covariance matrix is recursively computed with a rank-one update strategy, and the generalized eigenvalues and eigenvectors are approached with the projection approximation subspace tracking method. Numerical simulations show that the proposed recursive method is able to achieve satisfactory performance with significantly lower complexity as compared to the batch algorithm. Yingke Zhao, Jie Chen 0022, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2019 | Differential Kronecker Product BeamformingabstractDifferential beamformers have attracted much interest over the past few decades. In this paper, we introduce differential Kronecker product beamformers that exploit the structure of the steering vector to perform beamforming differently from the well-known and studied conventional approach. We consider a class of microphone arrays that enable to decompose the steering vector as a Kronecker product of two steering vectors of smaller virtual arrays. In the proposed approach, instead of directly designing the differential beamformer, we break it down following the decomposition of the steering vector, and show how to derive differential beamformers using the Kronecker product formulation. As demonstrated, the Kronecker product decomposition facilitates further flexibility in the design of differential beamformers and in the tradeoff control between the directivity factor and the white noise gain. Israel Cohen, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | On the Design of Target Beampatterns for Differential Microphone ArraysabstractDifferential microphone arrays (DMAs) have many interesting properties and have been widely used in acoustic, audio, and speech applications. A critical part of a DMA is the differential beamformer, which is generally designed in two important steps: 1) specifying a target beampattern based on what differential sound pressure field the DMA is expected to respond to and 2) designing the differential beamforming filter so that the resulting beampattern matches the target one. Most efforts in the study of DMAs so far have focused on the second step while choosing one of the limited patterns available in the literature as the target beampattern. Since it governs how the array performs, how to design the target beampattern is an important problem, which this paper addresses. The major contributions of this paper consists of the following four aspects. First, a positive superposition theorem is presented, which shows that the linear combination of effective beampatterns with non-negative coefficients is always an effective beampattern. Second, we propose a general approach to the design of target DMA beampatterns based on the positive superposition theorem. Third, an overview of the classical target beampatterns is provided and discussion is made on how to form effective base patterns. Fourth, we show that the smallest first null of a DMA is π(2N) with N being the DMA order, which provides the rule of setting nulls in practice. Finally, with examples, we show that with the use of the alternating-direction-method-of-multipliers algorithm, the proposed approach is able to generate useful DMA target beampatterns. Chao Pan 0001, Jingdong Chen, Jacob Benesty, Guangming Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | On Robust and High Directive Beamforming With Small-Spacing Microphone Arrays for Scattered SourcesabstractThis paper is devoted to beamforming with small-spacing microphone arrays for processing broadband and scattered acoustic sources. It presents a maximum diffuse noise gain (MDNG) beamformer in this context using the joint diagonalization technique, which is effective in suppressing diffuse and directional noise, but at a price of low white noise gain (WNG). We also introduce a maximum WNG (MWNG) beamformer, which is robust to the array imperfections, but paying a price of sacrificing the diffuse noise gain (DNG). To make a tradeoff between WNG and DNG so that the beamformer, on the one hand, can achieve high directivity and, on the other hand, is robust to implement, we propose a generalized MDNG beamformer, which includes both the MDNG and MWNG beamformers as particular cases. Simulations are conducted to illustrate the properties and advantages of the proposed beamformers. Xianghui Wang, Israel Cohen, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | On the Design of Robust Steerable Frequency-Invariant Beampatterns with Concentric Circular Microphone ArraysabstractThis paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs). We develop a beamforming algorithm based on an optimal approximation of the beamformer's beampattern with the Jacobi-Anger expansion. In comparison with the existing frequency-invariant beamformers with either circular microphone arrays (CMAs) or CCMAs, the developed algorithm offers the following advantages: 1) it can mitigate the deep-null problem encountered in CMAs and therefore has a consistent directivity factor over the frequency range of speech signals; 2) it is more flexible in terms of steering flexibility and the resulting beampattern can be steered to any direction; and 3) it does not require the microphones in different rings of the CCMA to be aligned, which is very useful in practice, particularly when microphone arrays with small and compact apertures have to be used. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2018 | Adaptive Parameters Adjustment for Group Reweighted Zero-Attracting LMSabstractInternational audience Danqi Jin, Jie Chen 0022, Cédric Richard, Jingdong Chen |
ICASSP | 4 |
| 2018 | On Speech Enhancement Using Microphone Arrays in the Presence of Co-Directional InterferenceabstractBeamforming using microphone arrays has been widely used for enhancing speech signals of interest and suppressing noise and interference in a wide range of applications. In order to make it work, beamforming generally assumes that the speech source of interest and the interference source are incident to the array from different directions. In this paper, we study the case where both the speech and interference sources come from the same direction. A linearly constrained minimum variance (LCMV) beamformer is derived in this scenario based on the so-called widely linear (WL) estimation framework in the frequency domain. We analyze this beamformer and show how its performance depends on the second-order non-circularity of the desired speech and interference sources. Xin Leng, Jingdong Chen, Jacob Benesty, Israel Cohen |
ICASSP | 2 |
| 2018 | A Single-Channel Noise Reduction Filtering/Smoothing Technique in the Time DomainabstractIn this paper, we present a single-channel smoothing-and-filtering technique for noise reduction in the time domain. Unlike traditional noise reduction methods, which directly apply a noise reduction filter to the noisy signal, the developed technique achieves noise reduction in two steps. It first applies a time smoothing window to the noisy signal, which, on the one hand, can help reduce high frequency noise and, on the other hand, can help leverage the correlation between successive signal samples. A noise reduction filter is then applied to the smoothed noisy signal to estimate the speech signal of interest. Three optimal and suboptimal noise reduction filters are derived, including the Wiener, maximum signal-to-noise-ratio (SNR), and tradeoff filters. Simulation results reveal that the developed method can produce better noise reduction performance, i.e., higher gains in the perceptual-evaluation-of-speech-quality (PESQ) score, than the traditional methods without smoothing. Ningning Pan, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2018 | Cosine Metric Learning for Speaker Verification in the I-vector Space
Zhongxin Bai, Xiao-Lei Zhang 0001, Jingdong Chen |
INTERSPEECH | 3 |
| 2018 | Model-driven online parameter adjustment for zero-attracting LMS
Danqi Jin, Jie Chen 0022, Cédric Richard, Jingdong Chen |
Signal Process. | 4 |
| 2018 | Noise Robust Frequency-Domain Adaptive Blind Multichannel Identification With ℓp-Norm ConstraintabstractBlind multichannel identification is a challenging problem in many domains. The normalized multichannel frequency-domain least-mean-square (NMCFLMS) algorithm was developed to blindly identify a single-input multiple-output acoustic system, which can yield good performance in noise-free environments. However, the robustness of this algorithm to noise has been shown to be problematic. One way to improve the robustness is by applying a constraint on the spectral flatness of the channel impulse responses, which led to the development of the so-called robust normalized multichannel frequency-domain least-mean-square (RNMCFLMS) algorithm. This spectral flatness constraint, however, may not be always proper or reasonable in realistic acoustic environments. In this paper, we develop an ℓp-norm constraint based robust normalized multichannel frequency-domain least-mean-square (ℓp-RNMCFLMS) algorithm. The ℓp-norm constraint is introduced into the NMCFLMS algorithm to control the effect of different ℓp-norm penalties on the adaptive filter for the impulse responses with different degrees of sparseness. Numerical and realistic experiments justify the effectiveness of the proposed ℓp-RNMCFLMS algorithm. Hongsen He, Jingdong Chen, Jacob Benesty, Tao Yang 0039 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Insights Into Frequency-Invariant Beamforming With Concentric Circular Microphone ArraysabstractThis paper studies the problem of frequency-invariant beamforming with concentric circular microphone arrays (CCMAs) and presents an approach to the design of frequency-invariant and symmetric beampatterns. We first apply the Jacobi-Anger expansion to each ring of the CCMA to approximate the beampattern. The beamformer is then designed by using all the expansions from different rings. In comparison with the existing work in the literature where a Jacobi-Anger expansion of the same order is applied to different rings, here in this contribution the order of the Jacobi-Anger expansion at a ring is related to its number of sensors and, as a result, the expansion order at different rings may be different. The developed approach is rather general. It is not only able to mitigate the deep nulls problem in the directivity factor and the white noise gain, that is common to circular microphone arrays (CMAs), and improve the steering flexibility, but is also flexible to use in practice where a smaller ring can have less microphones than a larger one. We discuss the conditions for the design ofNth-order symmetric beampatterns and examples of frequency-invariant beampatterns with commonly used array geometries such as CMAs, CMAs with a sensor at the center, and CCMAs. We show the advantage of adding one microphone at the center of either a CMA or a CCMA, i.e., circumventing the deep nulls problem caused by the 0th-order Bessel function. Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Robust multichannel TDOA estimation for speaker localization using the impulsive characteristics of speech spectrumabstractTime delay estimation (TDE) plays an important role in localizing and tracking radiating acoustic sources. Although many efforts have been devoted to this problem in the literature, the robustness of TDE with respect to noise and reverberation remains a great challenge for practical systems. In this paper, we investigate the TDE problem in acoustic single-input/multiple-output (SIMO) systems in reverberant and noisy environments. We first define a Cauchy estimator in the frequency domain, which is robust in dealing with speech as the SIMO system's excitation. This robust estimator is then used to construct a cost function, from which a robust multichannel frequency-domain adaptive filter is deduced. This adaptive algorithm is subsequently employed to blindly identify the acoustic impulse responses between the source and the microphones. Finally, the time difference of arrival is determined from the identified channel responses. Hongsen He, Jingdong Chen, Jacob Benesty, Yingyue Zhou, Tao Yang 0039 |
ICASSP | 2 |
| 2017 | Study of the frequency-domain multichannel noise reduction problem with the householder transformationabstractThis paper presents an approach to the multichannel noise reduction problem. It first transforms the multichannel noisy speech signals into the frequency domain. A Householder transformation is then constructed, which converts the multichannel coefficients in each frequency bin into two components: one dominated by speech and the other dominated by noise. A Wiener filter is subsequently formed to achieve an estimate of the noise in the speech dominated component from the noise dominated component. The enhanced speech is then obtained by subtracting the noise estimate from the speech dominated component. This approach consists of two critical steps: construction of the Householder transformation and formation of the noise reduction Wiener filter. If the source incidence angle is known a priori, the Householder transformation can be directly constructed using the steering vector and the optimal estimate of the signal of interest can then be obtained by applying the Wiener filter. If the source incidence angle is not known a priori, the Householder transformation can be constructed from a hypothesized incidence angle. Then, the optimal signal estimate is obtained by searching the maximum of the variance of the enhanced signal with the Wiener filter in the interested range of the incidence angle. Gongping Huang, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2017 | A minimum variance partially distortionless response filter for single-channel noise reductionabstractThis paper deals with the problem of single-channel noise reduction. Thanks to the eigenvalue decomposition, we arrange the eigenvalues of the speech correlation matrix in such a way that all the spectral mode signal-to-noise ratios (SNRs) of the noisy speech are ordered in a descending manner. By maintaining no speech distortion in the spectral modes with high input SNRs while allowing some degree of speech distortion in the modes with low input SNRs, we develop a minimum variance partially distortionless response (MVPDR) filter. We first formulate the problem and derive this filter within the general filtering framework. Then, the MVPDR filter is applied to the single-channel noise reduction problem in both the time and time-frequency domains. In comparison with the minimum variance distortionless response (MVDR) filter based on the subspace decomposition, the developed MVPDR filter can provide much more freedom for controlling the compromise between noise reduction and speech distortion to achieve higher speech quality. Simulations are conducted and preliminary results justify the advantages of the deduced MVPDR filter. Xianghui Wang, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2017 | On the Design of Frequency-Invariant Beampatterns With Uniform Circular Microphone ArraysabstractThis paper deals with two critical issues about uniform circular arrays (UCAs): frequency-invariant response and steering flexibility. It focuses on some optimal design of frequency-invariant beampatterns in any desired direction along the sensor plane. The major contributions are as follows. 1) We explain how to include the steering information in the desired directivity pattern. 2) We show that the optimal approximation of the beamformer's beampattern with a UCA from a least-squares error perspective is the Jacobi-Anger expansion. 3) We develop an approach to the design of any desired symmetric directivity pattern, where the deduced beampattern is almost frequency invariant and its main beam can be pointed to any wanted direction in the sensor plane. 4) With the proposed approach, we derive an explicit form of the white noise gain (WNG) and the directivity factor (DF), and explain clearly the white noise amplification problem at low frequencies and the DF degradation at high frequencies. The analysis also indicates that increasing the number of microphones can always improve the WNG. We show that the proposed method is a generalization of circular differential microphone arrays. The relationship between the proposed method and the so-called circular harmonics beamformers is also discussed. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | On time delay estimation based on multichannel spatiotemporal sparse linear predictionabstractNoise and reverberation can significantly affect the performance of time delay estimation (TDE) in room acoustic environments. The multichannel cross-correlation coefficient (MCCC) algorithm, which extends the traditional cross-correlation method from two to multiple channels, can exploit the spatial information among multiple microphones to improve the robustness of TDE with respect to environmental noise; but this algorithm is not robust to reverberation. The multichannel spatiotemporal prediction (MCSTP) algorithm uses both the spatial and temporal information provided by the array. This algorithm improves significantly the robustness of TDE with respect to reverberation; however, it is found sensitive to noise. In this paper, we develop a multichannel spatiotemporal sparse prediction (MCSTSP) algorithm for TDE. This algorithm obtains a good compromise between robustness of TDE to noise and that to reverberation through making a tradeoff between pre-whitening and non-prewhitening. This is achieved via adjusting a regularization parameter, which is solved by an augmented Lagrangian alternating direction method of multipliers (ADMM). The property of this developed algorithm is justified with numerical experiments in both noisy and reverberant environments. Hongsen He, Jingdong Chen, Jacob Benesty, Tao Yang 0039 |
ICASSP | 2 |
| 2016 | Subspace superdirective beamformers based on joint diagonalizationabstractAlthough they have been intensively studied and used in many applications due to their high directivity factor (DF), superdirective beamformers are sensitive to sensor noise and mismatch between sensors. This paper studies the problem of superdirective beamforming combined with the joint diagonalization method. We develop a subspace superdirective beamforming approach, which can achieve a good compromise between a high DF and white noise amplification. Simulations are performed to justify our theoretical analysis and demonstrate the good properties of this subspace superdirective beamforming approach. Changlei Li, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 4 |
| 2016 | A single-channel noise cancelation filter in the short-time-fourier-transform domainabstractThis paper develops a single-channel noise cancelation filter in the short-time Fourier transform (STFT) domain by combining the subspace method and the optimal filtering technique via joint diagonalization of the desired clean speech and noise signal correlation matrices. This filter is shown to be flexible in controlling the compromise between the output signal-to-noise ratio (oSNR) and the amount of speech distortion. Simulations are performed to justify the property of this filter. Xianghui Wang, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and MandarinabstractWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu |
ICML | 9 |
| 2016 | Single-channel noise reduction via semi-orthogonal transformations and reduced-rank filtering
Jacob Benesty, Jingdong Chen |
Speech Commun. | 3 |
| 2016 | Superdirective Beamforming Based on the Krylov MatrixabstractSuperdirective beamforming has attracted a significant amount of research interest in speech and audio applications, since it can maximize the directivity factor (DF) given an array geometry and, therefore, is efficient in dealing with signal acquisition in diffuse-like noise environments. However, this beamformer is very sensitive to sensor self-noise and mismatch among sensors, which considerably restricts its use in practical systems. This paper develops an approach to superdirective beamforming based on the Krylov matrix. We show that the columns of a proposed Krylov matrix, which span a chosen dimension of the whole space, are interesting beamformers; consequently, all different linear combinations of those columns lead to beamformers that have good properties. In particular, we develop the Krylov maximum white noise gain and Krylov maximum DF beamformers, which are obtained by maximizing the WNG and the DF, respectively. By properly choosing the dimension of the Krylov subspace, the developed beamformers that can make a compromise between reasonable values of the DF and white noise amplification. We also extend the basic idea to the design of the Krylov maximum front-to-back ratio, parametric superdirective, and parametric supercardioid beamformers. Gongping Huang, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Design of Directivity Patterns with a Unique Null of Maximum MultiplicityabstractDifferential beamforming is one of the most popular beamforming approaches, which has the great potential to form frequency-invariant directivity patterns. In this paper, we study the design of beampatterns with multiple nulls in the same direction, which is clearly different from the design of beampatterns with distinct nulls. Our contributions are as follows. First, we show how to constrain multiple nulls to the same direction and design the desired beampattern with both the traditional and robust approaches. Second, we derive an explicit form of the white noise gain (WNG) of the traditional approach as a function of the frequency, interelement spacing, and null direction, which shows that the cardioid is the optimal beampattern as far as the WNG is concerned. Third, we prove that the WNG improvement of the robust approach rarely depends on the null direction at low frequencies. Finally, considering the fact that the robust differential beamforming approach may produce a frequency-dependent beampattern while improving the WNG, we develop a weighted-norm approach that can make a good compromise between the robustness of differential beamforming with respect to white noise and the frequency-invariant beampattern. The performance of the developed approach is verified by simulations. Chao Pan 0001, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Reduced-Order Robust Superdirective Beamforming With Uniform Linear Microphone ArraysabstractSensor arrays for audio and speech signal acquisition are generally required to have frequency-invariant beampatterns to avoid adding spectral distortion to the broadband signals of interest. One way to obtain frequency-invariant beampatterns is via superdirective beamforming. However, traditional superdirective beamformers may cause significant white noise amplification (particularly at low frequencies), making them sensitive to uncorrelated white noise. To circumvent the problem of white noise amplification, a method was developed to find the superdirective beamforming filter with a constraint on the white noise gain (WNG), leading to the so-called WNG-constrained superdirective beamformer. But this method damages the frequency invariance of the beampattern. In this paper, we develop a flatness-constrained robust superdirective beamformer. We divide the overall beamformer into two subbeamformers, which are convolved together: one subbeamformer forms a lower order superdirective beampattern while the other attempts to improve the WNG. We show that this robust approach can improve the WNG while limiting the frequency dependency of the beampattern at the same time. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Investigation of a parametric gain approach to single-channel speech enhancementabstractThis paper investigates a parametric gain approach to single-channel noise reduction in the frequency domain. In comparison with the traditional parametric Wiener gain, the major novelty of this presented approach is that the parametric gain is formulated to estimate the noise by using the mean-squared error (MSE) between the noise and the noise estimate. The enhanced signal is then obtained by subtracting the noise estimate from the noisy observation signal. We show that this new method is more practical to implement and can produce better noise reduction performance as compared to the traditional parametric Wiener filtering techniques if the order of the parametric gain is not equal to 1. If the order is 1, the parametric gain is similar to the traditional Wiener gain. Simulation results are presented to illustrate the properties of this new approach. Gongping Huang, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2015 | Optimal single-channel noise reduction filtering matrices from the pearson correlation coefficient perspectiveabstractThis paper studies the problem of single-channel noise reduction in the time domain, where an estimate of a vector of the desired clean speech is achieved by filtering a frame of the noisy signal with a rectangular filtering matrix. The core issue with this problem formulation is then the estimation of the optimal filtering matrix. The squared Pearson correlation coefficient (SPCC) is used. We show that different optimal filtering matrices can be derived by maximizing or minimizing the SPCCs between different signals. For example, maximizing the SPCC between the enhanced signal and the filtered speech gives the reduced-rankWiener and minimum distortion (MD) filtering matrices while minimizing the SPCC gives the minimum noise (MN) and another reduced-rank Wiener filtering matrices. Simulation results are presented to illustrate the properties of these filtering matrices. Jiaolong Yu, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 4 |
| 2015 | Optimal design of directivity patterns for endfire linear microphone arraysabstractDirectivity pattern or beampattern is an important performance measure in all fixed beamformers. Given a microphone array, how to design the beamforming filter so that the resulting directivity pattern is close to the desired one is a critical issue. In this paper, we study the design of such patterns for endfire uniform linear microphone arrays. By considering the frequency-independent Chebyshev pattern as the desired one, we derive an optimal beamforming filter based on the minimization of the mean-squared error (MSE) under the distortionless constraint. It is shown that the proposed beamformer design can generate beampatterns that are very close to the desired ones and, the larger is the number of microphones, the better is the designed beampattern. Liheng Zhao, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2015 | Theoretical Analysis of Differential Microphone Array Beamforming and an Improved SolutionabstractDifferential microphone arrays (DMAs), which are responsive to the differential sound pressure field, have attracted much attention due to their properties of frequency-invariant beampatterns, small apertures, and potential of maximum directivity. Traditionally, DMAs are designed and implemented in a multistage (cascade) way, where a proper time delay is used in each stage to form a beampattern of interest. Recently, it was reported that DMAs can be designed by solving a linear system of equations formed from the information about the nulls of the desired beampattern. This paper deals with the problem of beamforming with linear DMAs. Its major contributions are as follows. 1) By using the spatial${\cal Z}$transform, we present some theoretical analysis of both the traditional cascade and new null-constrained DMA beamforming. It is shown that the cascade and null-constrained DMAs of the same order with the same number of sensors are theoretically identical. 2) We develop a two-stage approach to the study of the robust DMA beamformer, which is based on the principle of maximizing the white noise gain (WNG). The first-stage of this approach is in the structure of the traditional non-robust DMA while the second-stage filter is optimized for improving the WNG. 3) Using the two-stage approach, we show that the robust DMA beamformer may introduce extra nulls in the beampattern at high frequencies; particularly, it introduces$M - N - 1$extra nulls if the interelement spacing is equal to half of the wavelength, where$M$and$N$are the number of sensors and the DMA order, respectively. 4) We develop a method that can solve the extra-null problem while maximizing the WNG in robust DMA beamforming, i.e., a robust solution with a frequency-invariant beampattern. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | On the noisereduction performance of the MVDR beamformer innoisy and reverberant environmentsabstractThe minimum variance distortionless response (MVDR) beam-former has been widely studied for extraction of desired speech signals in noisy acoustic environments. The performance of this beam-former, however, depends on many factors such as the array geometry, the source incidence angle, the noise field characteristics, the reverberation conditions, etc. In this paper, we study the performance of the MVDR beamformer in different noise and reverberation conditions with a linear microphone array. Using the gain in signal-to-noise ratio (SNR) as the performance metric, we show that the optimal performance of the MVDR beamformer generally occurs when the source is in the endfire directions in different types of noise, which indicates that, as long as a linear array is used, we should configure it in such a way that the endfire direction is pointed to the desired source. Simulations in reverberant environments also verified this result, though the performance difference between end-fire and broadside directions reduces as the degree of reverberation increases. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
ICASSP | 2 |
| 2014 | Examples of optimal noise reduction filters derived from the squared Pearson correlation coefficientabstractThis paper studies the problem of single-channel noise reduction in the time domain. Based on some orthogonal decomposition developed recently and the squared Pearson correlation coefficient (SPCC), several noise reduction filters are derived. We will show that the optimization of the SPCC leads to the Wiener, minimum variance distortionless response (MVDR), minimum noise (MN), minimum uncorrelated speech and noise (MUSN), and linearly constrained minimum variance (LCMV) filters. We also compare the Wiener and MVDR filters derived from the SPCC to their counterparts derived from the mean-square error (MSE) criterion. Simulations are provided to illustrate the performance of all the deduced noise reduction filters. Jiaolong Yu, Jacob Benesty, Gongping Huang, Jingdong Chen |
ICASSP | 4 |
| 2014 | A family of maximum SNR filters for noise reductionabstractThis paper is devoted to the study and analysis of the maximum signal-to-noise ratio (SNR) filters for noise reduction both in the time and short-time Fourier transform (STFT) domains with one single microphone and multiple microphones. In the time domain, we show that the maximum SNR filters can significantly increase the SNR but at the expense of tremendous speech distortion. As a consequence, the speech quality improvement, measured by the perceptual evaluation of speech quality (PESQ) algorithm, is marginal if any, regardless of the number of microphones used. In the STFT domain, the maximum SNR filters are formulated by considering the interframe information in every frequency band. It is found that these filters not only improve the SNR, but also improve the speech quality significantly. As the number of input channels increases so is the gain in SNR as well as the speech quality. This demonstrates that the maximum SNR filters, particularly the multichannel ones, in the STFT domain may be of great practical value. Gongping Huang, Jacob Benesty, Tao Long 0004, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Multichannel Noise Reduction in the Karhunen-Loève Expansion DomainabstractThe noise reduction problem is traditionally approached in the time, frequency, or transform domain. Having a signal dependent transform has shown some advantages over the traditional signal independent transform. Recently, the single-channel noise reduction problem in the Karhunen-Loève expansion (KLE) domain has received special attention. In this paper, the noise reduction problem in the KLE domain is studied from a multichannel perspective. We present a new formulation of the problem, in which inter-channel and inter-mode correlations are optimally exploited. We derive different optimal noise reduction filters and present a set of useful performance measures within this framework. The performance of the different filters is then evaluated through experiments in which not only noise but also competing speech sources are present. It is shown that the proposed multichannel formulation is more robust to competing speech sources than the single-channel approach and that a better compromise between noise reduction and speech distortion can be obtained. Yesenia Lacouture-Parodi, Emanuël A. P. Habets, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Performance Study of the MVDR Beamformer as a Function of the Source Incidence AngleabstractLinear microphone arrays combined with the minimum variance distortionless response (MVDR) beamformer have been widely studied in various applications to acquire desired signals and reduce the unwanted noise. Most of the existing array systems assume that the desired sources are in the broadside direction. In this paper, we study and analyze the performance of the MVDR beamformer as a function of the source incidence angle. Using the signal-to-noise ratio (SNR) and beampattern as the criteria, we investigate its performance in four different scenarios: spatially white noise, diffuse noise, diffuse-plus-white noise, and point-source-plus-white noise. The results demonstrate that the optimal performance of the MVDR beamformer occurs when the source is in the endfire directions for diffuse noise and point-source noise while its SNR gain does not depend on the signal incidence angle in spatially white noise. This indicates that most current systems may not fully exploit the potential of the MVDR beamformer. This analysis does not only help us better understand this algorithm, but also helps us design better array systems for practical applications. Chao Pan 0001, Jingdong Chen, Jacob Benesty |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Design of Robust Differential Microphone ArraysabstractDifferential microphone arrays (DMAs), due to their small size and enhanced directivity, are quite promising in speech enhancement applications. However, it is well known that differential beamformers have the drawback of white noise amplification, which is a major issue in the processing of wideband signals such as speech. In this paper, we focus on the design of robust DMAs. Based on the Maclaurin’s series approximation and frequency-independent beampatterns, the robust first-, second-, and third-order DMAs are proposed by using more microphones than the order plus one, and the corresponding minimum-norm filters are derived. Compared to the traditional DMAs, the proposed designs are more robust with respect to white noise amplification while they are capable of achieving similar directional gains. Liheng Zhao, Jacob Benesty, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Multichannel acoustic echo suppressionabstractAcoustic echo suppression (AES) provides an attractive alternative to acoustic echo cancellation (AEC) techniques for full-duplex communication in low-complexity systems. However, so far AES techniques are commonly known to introduce significant distortions to the desired signal. Moreover, most traditional echo control techniques typically require accurately detecting the contribution of the near-end speaker to the microphone signal (“double talk”). The extension of AES techniques to the multichannel case usually assumes a symmetric system design which is often not fulfilled by typical scenarios. In this paper we propose a novel approach to multichannel acoustic echo suppression, which aims at extracting the near-end signal using a constraint for a distortionless output, without requiring a double-talk detector, or a symmetric system design. In addition to the above mentioned properties, the multichannel AES is also shown to overcome the known challenges in conventional multichannel acoustic echo control setups. Karim Helwani, Herbert Buchner, Jacob Benesty, Jingdong Chen |
ICASSP | 4 |
| 2013 | A study of the MVDR filter for acoustic echo suppressionabstractThis paper studies an echo suppression approach to reducing the undesired echoes that result from the acoustic coupling between a loudspeaker and a microphone in duplex voice communication. The approach consists of four basic steps. First, both the loudspeaker and microphone signals are partitioned into small overlapping frames. Second, each frame is transformed into the short-time Fourier transform (STFT) domain. Third, a minimum variance distortionless response (MVDR) filter is designed in each subband by explicitly using the interframe signal correlation. This MVDR filter is then used to estimate the echo signal and the obtained estimate is subsequently subtracted from the microphone signal. Finally, the time-domain processed signal is constructed using the overlap-add technique with the inverse STFT. Experiments are performed and the results demonstrate that this proposed method can achieve significant amount of echo suppression in practical room environments. Jacob Benesty, Jingdong Chen, Karim Helwani, Herbert Buchner |
ICASSP | 3 |
| 2013 | A Single-Channel MVDR Filter for Acoustic Echo SuppressionabstractAcoustic echo suppression techniques for full-duplex communication in low-complexity systems are commonly known to introduce distortion to the desired signal (i.e., near-end speech). Moreover, most traditional echo control techniques typically require accurately detecting the contribution of the near-end speaker to the microphone signal (“double talk”). In this letter, we propose a novel approach to acoustic echo suppression, which aims at extracting the near-end signal using a constraint for minimizing the distortion, and without requiring a double-talk detector. Karim Helwani, Herbert Buchner, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 4 |
| 2013 | On the Time-Domain Widely Linear LCMV Filter for Noise Reduction With a Stereo SystemabstractThis paper deals with the problem of noise reduction in stereo sound systems where the objective is not only to reduce noise, but also to preserve the spatial information of both the desired speech and noise sources so that the listener can still localize the speech and noise sources by listening to the enhanced binaural outputs. To achieve this objective, we use the widely linear (WL) framework developed previously and convert the problem of binaural noise reduction into one of monaural filtering with complex signals. We then present a way to decompose both the complex speech and noise signal vectors into two orthogonal components: one correlated and the other uncorrelated with the corresponding current signal sample. With this decomposition, the problem of noise reduction with preservation of the spatial information of speech and noise sources is formulated as an optimization problem with two constraints: one on the desired speech and the other on the preservation of the noise signal. We then derive a WL linearly constrained minimum variance (LCMV) filter, which can take advantage of the statistics and noncircularity of the complex speech signal to achieve noise reduction. In contrast to the WL Wiener and minimum variance distortionless response (MVDR) filters developed previously that can only preserve the characteristics and spatial information of the desired sound source, this new WL LCMV filter has the potential to reduce noise while preserving the characteristics and spatial information of both the desired and noise sources at the same time. Experimental results are provided to justify the claimed merits of the proposed WL LCMV filter. Jingdong Chen, Jacob Benesty |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Time Difference of Arrival Estimation Exploiting Multichannel Spatio-Temporal PredictionabstractTo localize sound sources in room acoustic environments, time differences of arrival (TDOA) between two or more microphone signals must be determined. This problem is often referred to as time delay estimation (TDE). The multichannel cross-correlation-coefficient (MCCC) algorithm, which is an extension of the traditional cross-correlation method from two- to multiple-channel cases, exploits spatial information among multiple microphones to improve the robustness of TDE. In this paper, we propose a multichannel spatio-temporal prediction (MCSTP) algorithm, which can be viewed as a generalization of the MCCC principle from using only spatial information to using both spatial and temporal information. A recursive version of this new algorithm is then developed, which can achieve similar performance as MCSTP, but is computationally more efficient. Experimental results in reverberant and noisy environments demonstrate the advantages of this new method for TDE. Hongsen He, Lifu Wu, Xiaojun Qiu, Jingdong Chen |
IEEE Trans. Speech Audio Process. | 5 |
| 2013 | A Class of Optimal Rectangular Filtering Matrices for Single-Channel Signal Enhancement in the Time DomainabstractIn this paper, we introduce a new class of optimal rectangular filtering matrices for single-channel speech enhancement. The new class of filters exploits the fact that the dimension of the signal subspace is lower than that of the full space. By doing this, extra degrees of freedom in the filters, that are otherwise reserved for preserving the signal subspace, can be used for achieving an improved output signal-to-noise ratio (SNR). Moreover, the filters allow for explicit control of the tradeoff between noise reduction and speech distortion via the chosen rank of the signal subspace. An interesting aspect is that the framework in which the filters are derived unifies the ideas of optimal filtering and subspace methods. A number of different optimal filter designs are derived in this framework, and the properties and performance of these are studied using both synthetic, periodic signals and real signals. The results show a number of interesting things. Firstly, they show how speech distortion can be traded for noise reduction and vice versa in a seamless manner. Moreover, the introduced filter designs are capable of achieving both the upper and lower bounds for the output SNR via the choice of a single parameter. Jesper Rindom Jensen, Jacob Benesty, Mads Græsbøll Christensen, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2012 | A multichannel widely linear approach to binaural noise reduction using an array of microphonesabstractThis paper deals with the problem of binaural noise reduction using an array of microphones. This is a very important problem in applications such as teleconferencing and hearing aids where there is a need to mitigate the noise effect from the noisy signals picked up by multiple microphones and produce two “clean” outputs. The mitigation of the noise should be made in such a way that no audible distortion is added to the two outputs (this is the same as in the single-channel case) and meanwhile the spatial information of the desired sound source should be preserved so that, after noise reduction, the listener will still be able to localize the sound source thanks to his/her binaural hearing mechanism. In this paper, we present a novel approach to this problem where we first form a number of complex input signals from the multiple and real microphone observations. We also merge the two expected real outputs into a complex output signal. The widely linear estimation theory is then used to derive optimal noise reduction filters that can achieve noise reduction while preserving the desired signal (speech) and its spatial information. With this new formulation, the Wiener and minimum variance distortionless response (MVDR) filters are derived. Experiments are provided to justify the effectiveness of these filters. Jacob Benesty, Jingdong Chen |
ICASSP | 2 |
| 2012 | Single-channel noise reduction in the STFT domain based on the bifrequency spectrumabstractThis paper studies the problem of noise reduction in the short-time Fourier transform (STFT) domain. Traditionally, the STFT coefficients in different frequency bands are assumed to be independent. This assumption holds when the signals are stationary and the fast Fourier transform(FFT) length is sufficiently large. In practice, however, speech is nonstationary and also the FFT length cannot be very large due to practical reasons. So, there always exists some correlation between STFT coefficients from neighboring frequency bands. An important question then arises: how the interband correlation can be used to optimize noise reduction performance? This paper addresses this issue. We discuss two solutions in the framework of the bifrequency spectrum. One considers the cross-correlation between all the frequency bands and the other takes into account only the cross-correlation between neighboring bands. While the former is optimal from a theoretical perspective, the latter is more practical as it is more immune to the error in correlation matrix estimation. Jingdong Chen, Jacob Benesty |
ICASSP | 1 |
| 2012 | Multi-microphone noise reduction using interchannel and interframe correlationsabstractMulti-microphone noise reduction methods often operate in the time-frequency domain in which a complex gain is applied to each time-frame and subband. These methods can achieve good noise reduction with little speech distortion by exploiting the fact that the desired signal is correlated across the channels. In the context of single-microphone noise reduction, it has been shown recently that the performance in terms of noise reduction and speech distortion can be improved by exploiting the correlation between subsequent time-frames, i.e., by exploiting the interframe correlation. In this paper, we exploit both interchannel and interframe correlations in the context of multi-microphone noise reduction. Now the interframe correlation is taken into account, i.e., a filter is applied in each subband and channel instead of just a gain. The results of our experimental study show that we can improve the fullband signal-to-noise ratios (SNRs) by using interchannel and interframe correlations when dealing with signals, such as speech, that exhibit a sufficiently large interframe correlation. Emanuël A. P. Habets, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2012 | Optimal rectangular filtering matrix for noise reduction in the time domainabstractIn this paper, we study the noise reduction problem in the time domain and present a frame-based method to decompose the clean speech vector into two orthogonal components: one correlated and the other uncorrelated with the current desired speech vector to be estimated. In comparison with the sample-based decomposition developed in the previous research that uses only forward prediction, this new decomposition exploits both the forward prediction and interpolation. Based on this new decomposition, we formulate different optimization cost functions and address the issue of how to design Wiener and minimum variance distortionless response (MVDR) filtering matrices by optimizing these new cost functions. We also discuss the relationship between the Wiener and MVDR filtering matrices and show that the MVDR filtering matrix can achieve noise reduction without adding speech distortion; but it reduces less noise than the Wiener filtering matrix. Compared with the sample-based algorithms developed in the previous study, the proposed frame-based algorithms can achieve better noise reduction performance. Furthermore, they are computationally more efficient, and therefore, more suitable for practical implementation. Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2011 | On single-channel noise reduction in the time domainabstractIn this paper, we revisit the noise-reduction problem in the time domain and present a way to decompose the filtered speech into two uncorrelated (orthogonal) components: the desired speech and the interference. Based on this new decomposition, we discuss how to form different optimization cost functions and address the issue of how to design different noise-reduction filters by optimizing these new cost functions. Particularly, we cover the design of the maximum signal-to-noise-ratio (SNR), the Wiener, the minimum variance distortionless response (MVDR), and the tradeoff filters. It is interesting that with this new decomposition, we can now design the MVDR filter that can achieve noise reduction without adding speech distortion in the single-channel case, which has never been seen before. We also demonstrate that the maximum SNR, Wiener, and tradeoff filters are identical to the MVDR filter up to a scaling factor. From a theoretical point of view, this scaling factor is not significant and should not affect the output SNR at any processing time. But from a practical viewpoint, the scaling factor can be time-varying due to the nonstationarity of the speech and possibly the noise and can cause discontinuity in the residual noise level, which is unpleasant to listen to. As a result, it is essential to have the scaling factor right from one processing sample (or frame) to another in order to avoid large distortions and for this reason, it is recommended to use the MVDR filter in speech enhancement applications. Jingdong Chen, Jacob Benesty, Yiteng Huang, Tomas Gänsler |
ICASSP | 1 |
| 2011 | Binaural Noise Reduction in the Time Domain With a Stereo SetupabstractBinaural noise reduction with a stereophonic (or simply stereo) setup has become a very important problem as stereo sound systems and devices are being more and more deployed in modern voice communications. This problem is very challenging since it requires not only the reduction of the noise at the stereo inputs, but also the preservation of the spatial information embodied in the two channels so that after noise reduction the listener can still localize the sound source from the binaural outputs. As a result, simply applying a traditional single-channel noise reduction technique to each channel individually may not work as the spatial effects may be destroyed. In this paper, we present a new formulation of the binaural noise reduction problem in stereo systems. We first form a complex signal from the stereo inputs with one channel being its real part and the other being its imaginary part. By doing so, the binaural noise reduction problem can be processed by a single-channel widely linear filter. The widely linear estimation theory is then used to derive optimal noise reduction filters that can fully take advantage of the noncircularity of the complex speech signal to achieve noise reduction while preserving the desired signal (speech) and spatial information. With this new formulation, the Wiener, minimum variance distortionless response (MVDR), maximum signal-to-noise ratio (SNR), and tradeoff filters are derived. Experiments are provided to justify the effectiveness of these filters. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | An Integrated Solution for Online Multichannel Noise Tracking and ReductionabstractNoise statistics estimation is a paramount issue in the design of reliable noise-reduction algorithms. Although significant efforts have been devoted to this problem in the literature, most developed methods so far have focused on the single-channel case. When multiple microphones are used, it is important that the data from all the sensors are optimally combined to achieve judicious updates of the noise statistics and the noise-reduction filter. This contribution is devoted to the development of a practical approach to multichannel noise tracking and reduction. We combine the multichannel speech presence probability (MC-SPP) that we proposed in an earlier contribution with an alternative formulation of the minima-controlled recursive averaging (MCRA) technique that we generalize from the single-channel to the multichannel case. To demonstrate the effectiveness of the proposed MC-SPP and multichannel noise estimator, we integrate them into three variants of the multichannel noise reduction Wiener filter. Experimental results show the advantages of the proposed solution. Mehrez Souden, Jingdong Chen, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Study of the widely linear Wiener filter for noise reductionabstractThis paper develops a new widely linear noise-reduction Wiener filter based on the variance and pseudo-variance of the short-time Fourier transform coefficients of speech signals. We show that this new noise-reduction filter has many interesting properties, including but not limited to: 1) it causes less speech distortion as compared to the classical noise-reduction Wiener filter; 2) its minimum mean-squared error (MSE) is smaller than that of the classical Wiener filter; 3) it can increase the subband signal-to-noise ratio (SNR), while the classical Wiener filter has no effect on the subband SNR for any given signal frame and subband. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP | 2 |
| 2010 | Analysis of the frequency-domain Wiener filter with the prediction gainabstractThis paper presents a theoretical analysis on the performance of the optimal noise-reduction filter in the frequency domain. Using the autoregressive (AR) model to model both the clean speech and noise, we build the relationship between the Wiener filter and the AR parameters of the clean speech and noise signals. We show that if noise is not predictable, the Wiener filter is mostly related to the AR parameters of the desired speech signal. On the contrary, if the desired signal is not predictable, the Wiener filter is then mostly related to the AR parameters of the noise signal. More importantly, we provide the bounds for noise reduction, speech distortion, and SNR improvement, and show that the performance of the Wiener filter in terms of SNR improvement and degree of noise reduction and speech distortion is closely related to the prediction gain of the desired speech and noise signals. Jingdong Chen, Jacob Benesty, Yiteng Huang |
ICASSP | 1 |
| 2010 | On widely linear Wiener and tradeoff filters for noise reduction
Jacob Benesty, Jingdong Chen, Yiteng Huang |
Speech Commun. | 2 |
| 2010 | A Widely Linear Distortionless Filter for Single-Channel Noise ReductionabstractTraditionally in the single-channel noise-reduction problem, speech distortion is inevitable since the desired signal is also filtered while filtering the noise. In fact, the more the noise is reduced, the more the speech distortion is added into the desired signal, as proved in the literature. So, if we require no speech distortion, we either end up with no noise reduction at all or have to use multiple sensors. In this paper, we attempt to apply the widely linear (WL) estimation theory to noise reduction. Unlike the traditional approaches that only filter the short-time Fourier transform (STFT) of the noisy signal, the method developed in this paper applies the noise-reduction filter to both the STFT of the noisy signal and its conjugate. With the constraint of no speech distortion, a WL distortionless filter is derived. We show that this new optimal filter can fully take advantage of the noncircularity property of speech signals to achieve up to 3-dB signal-to-noise-ratio (SNR) improvement without introducing any speech distortion, which can only be obtained with the traditional approaches if two or more microphones are used. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Signal Process. Lett. | 2 |
| 2010 | Gaussian Model-Based Multichannel Speech Presence ProbabilityabstractThe knowledge of the target speech presence probability in a mixture of signals captured by a speech communication system is of paramount importance in several applications including reliable noise reduction algorithms. In this correspondence, we establish a new expression for speech presence probability when an array of microphones with an arbitrary geometry is used. Our study is based on the assumption of the Gaussian statistical model for all signals and involves the noise and noisy data statistics only. In comparison with the single-channel case, the new proposed multichannel approach can significantly increase the detection accuracy. In particular, when the additive noise is spatially coherent, perfect speech presence detection is theoretically possible, while when the noise is spatially white, a coherent summation of speech components is performed to allow for enhanced speech presence probability estimation. Mehrez Souden, Jingdong Chen, Jacob Benesty, Sofiène Affes |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | On noise reduction in the Karhunen-Loève expansion domainabstractIn this paper, we study the noise-reduction problem in the Karhunen-Loève expansion domain. We develop two classes of optimal filters. The first class estimates a frame of speech by filtering the corresponding frame of the noisy speech. We will show that several well-known existing methods belong or are closely related to this category. The second class, which has not been studied before, obtains noise reduction by filtering not only the current frame, but also a number of previous consecutive frames of the noisy speech. We will discuss how to design the optimal noise-reduction filters in each class and demonstrate the properties of the deduced optimal filters. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP | 2 |
| 2009 | Using the Pearson correlation coefficient to develop an optimally weighted cross relation based blind SIMO identification algorithmabstractBlind SIMO identification is challenging when additive noise is strong and for ill-conditioned/acoustic SIMO systems. A weighted cross relation (CR) algorithm presumably can be robust to noise but there lacks a practical way to define the weights. In this paper, the Pearson correlation coefficient (PCC) is used to develop an optimally weighted CR algorithm, which is validated by simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP | 3 |
| 2009 | Noise Reduction Algorithms in a Generalized Transform DomainabstractNoise reduction for speech applications is often formulated as a digital filtering problem, where the clean speech estimate is obtained by passing the noisy speech through a linear filter/transform. With such a formulation, the core issue of noise reduction becomes how to design an optimal filter (based on the statistics of the speech and noise signals) that can significantly suppress noise without introducing perceptually noticeable speech distortion. The optimal filters can be designed either in the time or in a transform domain. The advantage of working in a transform space is that, if the transform is selected properly, the speech and noise signals may be better separated in that space, thereby enabling better filter estimation and noise reduction performance. Although many different transforms exist, most efforts in the field of noise reduction have been focused only on the Fourier and Karhunen-Loeve transforms. Even with these two, no formal study has been carried out to investigate which transform can outperform the other. In this paper, we reformulate the noise reduction problem into a more generalized transform domain. We will show some of the advantages of working in this generalized domain, such as 1) different transforms can be used to replace each other without any requirement to change the algorithm (optimal filter) formulation, and 2) it is easier to fairly compare different transforms for their noise reduction performance. We will also address how to design different optimal and suboptimal filters in such a generalized transform domain. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Study of the Noise-Reduction Problem in the Karhunen-LoÈve Expansion DomainabstractNoise reduction, which aims at estimating a clean speech from a noisy observation, has long been an active research area. The standard approach to this problem is to obtain the clean speech estimate by linearly filtering the noisy signal. The core issue, then, becomes how to design an optimal linear filter that can significantly suppress noise without introducing perceptually noticeable speech distortion. Traditionally, the optimal noise-reduction filters are formulated in either the time or the frequency domains. This paper studies the problem in the Karhunen–LoÈve expansion domain. We develop two classes of optimal filters. The first class achieves a frame of speech estimate by filtering the corresponding frame of the noisy speech. We will show that many existing methods such as the widely used Wiener filter and subspace technique are closely related to this category. The second class obtains noise reduction by filtering not only the current frame, but also a number of previous consecutive frames of the noisy speech. We will discuss how to design the optimal noise-reduction filters in each class and demonstrate, through both theoretical analysis and experiments, the properties of the deduced optimal filters. Jingdong Chen, Jacob Benesty, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | A minimum speech distortion multichannel algorithm for noise reductionabstractNoise reduction using multiple microphones remains a challenging and crucial research problem. This paper presents a new multichannel noise-reduction algorithm based on spatio-temporal prediction. Unlike many multichannel techniques that attempt to achieve both speech dereverberation and noise reduction at the same time, this new approach puts aside speech dereverberation and formulates the problem as one of estimating the speech component received at one microphone using the observations from all the available microphones. In comparison with the existing techniques such as beamforming, this new multichannel approach has many appealing properties: it does not require the knowledge of the source location or the channel impulse responses; the multiple microphones do not have to be arranged into a specific array geometry; it works the same for both the far-field and near-field cases; and most importantly, it can produce very good noise reduction with minimum speech distortion in real acoustic environments. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP | 2 |
| 2008 | Design of Steerable Linear Differential Microphone Arrays With Omnidirectional and Bidirectional SensorsabstractThis paper is dedicated to the design of fully steerable linear differential microphone arrays (LDMAs). We analyze the steerable ideal spatial responses and explain why conventional LDMAs consisting of only omnidirectional microphones have limited steering ability. In order to circumvent this limitation, we suggest to use both omnidirectional and bidirectional (with a dipole shaped directivity pattern) microphones. We discuss the minimum numbers of omnidirectional and bidirectional sensors required for achieving steerable spatial responses and present a method to design fully steerable differential beamformers with LDMAs through the Jacobi-Anger series expansion. Simulations validate the presented technique and the steering flexibility of the designed LDMAs. Xueqin Luo, Jilu Jin, Gongping Huang, Jingdong Chen, Jacob Benesty |
IEEE Signal Process. Lett. | 4 |
| 2008 | On the Importance of the Pearson Correlation Coefficient in Noise ReductionabstractNoise reduction, which aims at estimating a clean speech from noisy observations, has attracted a considerable amount of research and engineering attention over the past few decades. In the single-channel scenario, an estimate of the clean speech can be obtained by passing the noisy signal picked up by the microphone through a linear filter/transformation. The core issue, then, is how to find an optimal filter/transformation such that, after the filtering process, the signal-to-noise ratio (SNR) is improved but the desired speech signal is not noticeably distorted. Most of the existing optimal filters (such as the Wiener filter and subspace transformation) are formulated from the mean-square error (MSE) criterion. However, with the MSE formulation, many desired properties of the optimal noise-reduction filters such as the SNR behavior cannot be seen. In this paper, we present a new criterion based on the Pearson correlation coefficient (PCC). We show that in the context of noise reduction the squared PCC (SPCC) has many appealing properties and can be used as an optimization cost function to derive many optimal and suboptimal noise-reduction filters. The clear advantage of using the SPCC over the MSE is that the noise-reduction performance (in terms of the SNR improvement and speech distortion) of the resulting optimal filters can be easily analyzed. This shows that, as far as noise reduction is concerned, the SPCC-based cost function serves as a more natural criterion to optimize as compared to the MSE. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | A Minimum Distortion Noise Reduction Algorithm With Multiple MicrophonesabstractThe problem of noise reduction using multiple microphones has long been an active area of research. Over the past few decades, most efforts have been devoted to beamforming techniques, which aim at recovering the desired source signal from the outputs of an array of microphones. In order to work reasonably well in reverberant environments, this approach often requires such knowledge as the direction of arrival (DOA) or even the room impulse responses, which are difficult to acquire reliably in practice. In addition, beamforming has to compromise its noise reduction performance in order to achieve speech dereverberation at the same time. This paper presents a new multichannel algorithm for noise reduction, which formulates the problem as one of estimating the speech component observed at one microphone using the observations from all the available microphones. This new approach explicitly uses the idea of spatial–temporal prediction and achieves noise reduction in two steps. The first step is to determine a set of inter-sensor optimal spatial–temporal prediction transformations. These transformations are then exploited in the second step to form an optimal noise-reduction filter. In comparison with traditional beamforming techniques, this new method has many appealing properties: it does not require DOA information or any knowledge of either the reverberation condition or the channel impulse responses; the multiple microphones do not have to be arranged into a specific array geometry; it works the same for both the far-field and near-field cases; and, most importantly, it can produce very good and robust noise reduction with minimum speech distortion in practical environments. Furthermore, with this new approach, it is possible to apply postprocessing filtering for additional noise reduction when a specified level of speech distortion is allowed. Jingdong Chen, Jacob Benesty, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Analysis and Comparison of Multichannel Noise Reduction Methods in a Common FrameworkabstractNoise reduction for speech enhancement is a useful technique, but in general it is a challenging problem. While a single-channel algorithm is easy to use in practice, it inevitably introduces speech distortion to the desired speech signal while reducing noise. Today, the explosive growth in computational power and the continuous drop in the cost and size of acoustic electric transducers are driving the interest of employing multiple microphones in speech processing systems. This opens new opportunities for noise reduction. In this paper, we present an analysis of three multichannel noise reduction algorithms, namely Wiener filter, subspace, and spatial-temporal prediction, in a common framework. We intend to investigate whether it is possible for the multichannel noise reduction algorithms to reduce noise without speech distortion. Finally, we justify what we learn via theoretical analyses by simulations using real impulse responses measured in the varechoic chamber at Bell Labs. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | On Recursive and Fast Recursive Computation of the Capon SpectrumabstractThe Capon spectrum, which is known to have better resolution than the periodogram, has been widely used in various applications. Normally, the Capon spectrum is estimated through the direct computation of the inverse of the data correlation (or covariance) matrix. This so-called direct inverse approach is, however, computationally very expensive due to the high computational cost involved in the matrix inversion. This paper deals with fast and efficient algorithms in computing the Capon spectrum. Inspired from the recursive idea established in the area of adaptive signal processing, we first derive a recursive Capon algorithm. This new algorithm does not require an explicit matrix inversion, and hence is more efficient to implement than the direct inverse method. We then develop a fast version of the recursive algorithm, which can further reduce the complexity of the recursive one by an order of magnitude. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP (3) | 2 |
| 2007 | An Acoustic MIMO Framework for Analyzing Microphone-Array BeamformingabstractAlthough a significant amount of research attention has been devoted to microphone-array beamforming, the performance of all the developed algorithms in practical acoustic environments is still far from meeting our expectation. So further research efforts on this topic are indispensable. In this paper, we treat a microphone array as a multiple-input multiple-output (MIMO) system and develop a general framework for analyzing performance of beamforming algorithms based on the acoustic MIMO channel impulse responses. Under this framework, we study the bounds for the length of beamforming filter, which in turn shows the performance bounds of beamforming in terms of speech dereverberation and interference suppression. We also discuss the intrinsic relationships among different classical beamforming techniques and explain, from the channel condition point of view, what the prerequisites have to be fulfilled in order for those techniques to work. Jingdong Chen, Jacob Benesty, Yiteng Huang |
ICASSP (1) | 1 |
| 2007 | Laplace Entropy and its Application to Time Delay Estimation for Speech SignalsabstractTime delay estimation (TDE) is a basic technique for numerous applications where there is a need to localize and track a radiating source. It is particularly challenging in the presence of noise and reverberation, and when the source signal is speech which is inherently nonstationary and random. The most important TDE algorithms for two sensors are based on the generalized cross-correlation (GCC) method. These algorithms perform reasonably well when reverberation or noise is not too high. In an earlier study of the authors, a more sophisticated approach was proposed. It employs more sensors and takes advantage of their delay redundancy to improve the precision of the TDOA (time difference of arrival) estimate between the first two sensors. The approach is based on the multichannel cross-correlation coefficient (MCCC) and was found more robust to noise and reverberation. In this paper, we show that this approach can also be developed on a basis of joint entropy. For Gaussian signals, we show that, in the search of the TDOA estimate, maximizing MCCC is equivalent to minimizing joint entropy. But with the generalization of the idea to non-Gaussian speech signals, the joint entropy based new multichannel TDE algorithm manifests a potential to outperform the MCCC-based method. Since there is no rigorous mathematical formula for speech entropy, we use the assumption that speech can be plausibly modeled by a Laplace distribution and develop a practical approximation of Laplace entropy for TDE of speech signals. The performance of the proposed new algorithm is investigated via simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (1) | 3 |
| 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learningabstractSpeech dereverberation remains an open problem after more than three decades of research. The most challenging step in speech dereverberation is blind chan- nel identification (BCI). Although many BCI approaches have been developed, their performance is still far from satisfactory for practical applications. The main difficulty in BCI lies in finding an appropriate acoustic model, which not only can effectively resolve solution degeneracies due to the lack of knowledge of the source, but also robustly models real acoustic environments. This paper proposes a sparse acoustic room impulse response (RIR) model for BCI, that is, an acous- tic RIR can be modeled by a sparse FIR filter. Under this model, we show how to formulate the BCI of a single-input multiple-output (SIMO) system into a l1- norm regularized least squares (LS) problem, which is convex and can be solved efficiently with guaranteed global convergence. The sparseness of solutions is controlled by l1-norm regularization parameters. We propose a sparse learning scheme that infers the optimal l1-norm regularization parameters directly from microphone observations under a Bayesian framework. Our results show that the proposed approach is effective and robust, and it yields source estimates in real acoustic environments with high fidelity to anechoic chamber measurements. Yuanqing Lin, Jingdong Chen, Youngmoo E. Kim, Daniel D. Lee |
NIPS | 2 |
| 2007 | On the optimal linear filtering techniques for noise reduction
Jingdong Chen, Jacob Benesty, Yiteng Huang |
Speech Commun. | 1 |
| 2007 | Time Delay Estimation via Minimum EntropyabstractTime delay estimation (TDE) is a basic technique for numerous applications where there is a need to localize and track a radiating source. The most important TDE algorithms for two sensors are based on the generalized cross-correlation (GCC) method. These algorithms perform reasonably well when reverberation or noise is not too high. In an earlier study by the authors, a more sophisticated approach was proposed. It employs more sensors and takes advantage of their delay redundancy to improve the precision of the time difference of arrival (TDOA) estimate between the first two sensors. The approach is based on the multichannel cross-correlation coefficient (MCCC) and was found more robust to noise and reverberation. In this letter, we show that this approach can also be developed on a basis of joint entropy. For Gaussian signals, we show that, in the search of the TDOA estimate, maximizing MCCC is equivalent to minimizing joint entropy. However, with the generalization of the idea to non-Gaussian signals (e.g., speech), the joint entropy-based new TDE algorithm manifests a potential to outperform the MCCC-based method Jacob Benesty, Yiteng Huang, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2007 | On Crosstalk Cancellation and Equalization With Multiple Loudspeakers for 3-D Sound ReproductionabstractPeople prefer to be able to enjoy spatial audio without wearing a headphone. Such a tethered device is anyway inconvenient and undesirable, if not cumbersome. Alternatively, 3D sound can be delivered to a listener with loudspeakers. However, crosstalk arises, and the rendered binaural signals are distorted by room reverberation when arriving at the listener's two ears, which lead to the need for a crosstalk cancellation and equalization (CTCE) system. Classical CTCE systems employ only two loudspeakers, and their performance is usually unsatisfactory in practice. While the idea of using more loudspeakers has been investigated, it was never shown why using more loudspeakers is theoretically more advantageous for CTCE. In this letter, we will study this problem and demonstrate that with two loudspeakers, only a least-squares (LS) solution can be obtained, while using multiple loudspeakers, we have more options: either an LS solution or an exact solution for perfect CTCE. These findings are justified by simulations using real impulse responses measured in the varechoic chamber at Bell Labs. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2007 | On Microphone-Array Beamforming From a MIMO Acoustic Signal Processing PerspectiveabstractAlthough many microphone-array beamforming algorithms have been developed over the past few decades, most such algorithms so far can only offer limited performance in practical acoustic environments. The reason behind this has not been fully understood and further research on this matter is indispensable. In this paper, we treat a microphone array as a multiple-input multiple-output (MIMO) system and study its signal-enhancement performance. Our major contribution is fourfold. First, we develop a general framework for analyzing performance of beamforming algorithms based on the acoustic MIMO channel impulse responses. Second, we study the bounds for the length of the beamforming filter, which in turn shows the performance bounds of beamforming in terms of speech dereverberation and interference suppression. Third, we address the connection between beamforming and the multiple-input/output inverse theorem (MINT). Finally, we discuss the intrinsic relationships among different classical beamforming techniques and explain, from the channel condition perspective, what the prerequisites are for those techniques to work. Jacob Benesty, Jingdong Chen, Yiteng Huang, Jacek Dmochowski |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Estimation of the Coherence Function with the MVDR ApproachabstractThe minimum variance distortionless response (MVDR), originally developed by Capon for frequency-wavenumber analysis, is a very well established method in array processing. It is also used in spectral estimation. The aim of this paper is to show how the MVDR method can be used to estimate the magnitude squared coherence (MSC) function, which is very useful in so many applications but so few methods exist to estimate it. Simulations show that our algorithm gives much more reliable results than the one based on the popular Welch's method. Jacob Benesty, Jingdong Chen, Yiteng Huang |
ICASSP (3) | 2 |
| 2006 | Speech Acquisition and Enhancement in a Reverberant, Cocktail-Party-Like EnvironmentabstractDeveloping a successful multi-microphone speech acquisition system in a reverberant, cocktail-party-like environment is a very challenging problem since both interfering sources and reverberation need to be well controlled. In this paper, we propose an algorithm based on blind SIMO identification. We first blindly identify the channels from the interfering sources to all the microphones. Then we extract the speech signal of interest. Finally speech dereverberation is performed using the MINT method. Simulations with acoustic impulse responses measured in the varechoic chamber at Bell Labs are carried out to verify the proposed algorithm. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (5) | 3 |
| 2006 | Identification of acoustic MIMO systems: Challenges and opportunities
Yiteng Huang, Jacob Benesty, Jingdong Chen |
Signal Process. | 3 |
| 2006 | New insights into the noise reduction Wiener filterabstractThe problem of noise reduction has attracted a considerable amount of research attention over the past several decades. Among the numerous techniques that were developed, the optimal Wiener filter can be considered as one of the most fundamental noise reduction approaches, which has been delineated in different forms and adopted in various applications. Although it is not a secret that the Wiener filter may cause some detrimental effects to the speech signal (appreciable or even significant degradation in quality or intelligibility), few efforts have been reported to show the inherent relationship between noise reduction and speech distortion. By defining a speech-distortion index to measure the degree to which the speech signal is deformed and two noise-reduction factors to quantify the amount of noise being attenuated, this paper studies the quantitative performance behavior of the Wiener filter in the context of noise reduction. We show that in the single-channel case the a posteriori signal-to-noise ratio (SNR) (defined after the Wiener filter) is greater than or equal to the a priori SNR (defined before the Wiener filter), indicating that the Wiener filter is always able to achieve noise reduction. However, the amount of noise reduction is in general proportional to the amount of speech degradation. This may seem discouraging as we always expect an algorithm to have maximal noise reduction without much speech distortion. Fortunately, we show that speech distortion can be better managed in three different ways. If we have some a priori knowledge (such as the linear prediction coefficients) of the clean speech signal, this a priori knowledge can be exploited to achieve noise reduction while maintaining a low level of speech distortion. When no a priori knowledge is available, we can still achieve a better control of noise reduction and speech distortion by properly manipulating the Wiener filter, resulting in a suboptimal Wiener filter. In case that we have multiple microphone sensors, the multiple observations of the speech signal can be used to reduce noise with less or even no speech distortion. Jingdong Chen, Jacob Benesty, Yiteng Huang, Simon Doclo |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Time delay estimation via multichannel cross-correlation [audio signal processing applications]abstractTime delay estimation (TDE) in a reverberant acoustical environment is a very challenging and difficult problem. This paper tackles the problem by exploiting the redundant information provided by multiple microphone sensors. To do so, the multichannel crosscorrelation coefficient (MCCC) is re-derived, in a new way, to connect it to the well-known linear interpolation technique. Some interesting properties and bounds of MCCC are discussed, and a recursive algorithm is then introduced so that MCCC can be estimated and updated efficiently when new data snapshots are available. We then apply the MCCC to the TDE problem, resulting a multichannel cross-correlation algorithm that can be treated as a natural generalization of the generalized cross-correlation (GCC) TDE method to the multichannel case. It is shown that this method can take advantage of the redundancy provided by multiple microphone sensors to improve TDE against both reverberation and noise. Jingdong Chen, Yiteng Huang, Jacob Benesty |
ICASSP (3) | 1 |
| 2005 | Adaptive blind SIMO identification: derivation of an optimal step size for the unconstrained multichannel LMS algorithmabstractAdaptive algorithms for blindly identifying SIMO systems are appealing because of their computational efficiency and capability of continuously tracking a time-varying system. Adaptive multichannel LMS (MCLMS) algorithms (with and without the unit-norm constraint) are analyzed and the optimal step size is derived. A simple yet effective variable step-size unconstrained MCLMS algorithm is proposed and its performance is evaluated with simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (3) | 3 |
| 2005 | A generalized MVDR spectrumabstractThe minimum variance distortionless response (MVDR) approach is very popular in array processing. It is also employed in spectral estimation where the Fourier matrix is used in the optimization process. First, we give a general form of the MVDR where any unitary matrix can be used to estimate the spectrum. Second and most importantly, we show how the MVDR method can be used to estimate the magnitude squared coherence function, which is very useful in so many applications but so few methods exist to estimate it. Simulations show that our algorithm gives much more reliable results than the one based on the popular Welch's method. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Signal Process. Lett. | 2 |
| 2005 | Optimal step size of the adaptive multichannel LMS algorithm for blind SIMO identificationabstractAdaptive algorithms for blindly identifying single-input multiple-output (SIMO) systems are appealing because of their computational efficiency and capability of continuously tracking a time-varying system. Adaptive multichannel least-mean-square (MCLMS) algorithms (with and without the unit-norm constraint) are analyzed, and the optimal step size is derived. A simple yet effective variable step-size MCLMS algorithm is proposed, and its performance is evaluated with simulations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 3 |
| 2005 | Suppressing Acoustic Echo in a Spectral Envelope SpaceabstractFull-duplex hands-free telecommunication systems employ an acoustic echo canceler (AEC) to remove the undesired echoes that result from the coupling between a loudspeaker and a microphone. Traditionally, the removal is achieved by modeling the echo path impulse response with an adaptive finite impulse response (FIR) filter and subtracting an echo estimate from the microphone signal. It is not uncommon that an adaptive filter with a length of 50-300 ms needs to be considered, which makes an AEC highly computationally expensive. In this paper, we propose an echo suppression algorithm to eliminate the echo effect. Instead of identifying the echo path impulse response, the proposed method estimates the spectral envelope of the echo signal. The suppression is done by spectral modification-a technique originally proposed for noise reduction. It is shown that this new approach has several advantages over the traditional AEC. Properties of human auditory perception are considered, by estimating spectral envelopes according to the frequency selectivity of the auditory system, resulting in improved perceptual quality. A conventional AEC is often combined with a post-processor to reduce the residual echoes due to minor echo path changes. It is shown that the proposed algorithm is insensitive to such changes. Therefore, no post-processor is necessary. Furthermore, the new scheme is computationally much more efficient than a conventional AEC. Christof Faller, Jingdong Chen |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | A Blind Channel Identification-Based Two-Stage Approach to Separation and Dereverberation of Speech Signals in a Reverberant EnvironmentabstractBlind separation of independent speech sources from their convolutive mixtures in a reverberant acoustic environment is a difficult problem and the state-of-the-art blind source separation techniques are still unsatisfactory. The challenge lies in the coexistence of spatial interference from competing sources and temporal echoes due to room reverberation in the observed mixtures. Focusing only on optimizing the signal-to-interference ratio is inadequate for most if not all speech processing systems. In this paper, we deduce that spatial interference and temporal echoes can be separated and an M/spl times/N MIMO system will be converted into M SIMO systems that are free of spatial interference. Furthermore we show that the channel matrices of these SIMO systems are irreducible if the channels from the same source in the MIMO system do not share common zeros. Thereafter we can apply the Bezout theorem to remove reverberation in those SIMO systems. Such a two-stage procedure leads to a novel sequential source separation and speech dereverberation algorithm based on blind multichannel identification. Simulations with measurements obtained in the varechoic chamber at Bell Labs demonstrate the success and robustness of the proposed algorithm in highly reverberant acoustic environments. Yiteng Huang, Jacob Benesty, Jingdong Chen |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | An exponentiated gradient adaptive algorithm for blind identification of sparse SIMO systemsabstractSparse impulse responses are encountered in many acoustic and wireless channels. Recently, a class of exponentiated gradient (EG) algorithms has been proposed. One of the algorithms belonging to this class, the so-called EG/spl plusmn/ algorithm, converges and tracks much better than the classical stochastic gradient, or LMS, algorithm for sparse impulse responses. We apply this technique to blind identification of a sparse SIMO system and develop the multichannel EG/spl plusmn/ algorithm. A simple experiment demonstrates its advantage in convergence compared to the MCLMS algorithm. Jacob Benesty, Yiteng Huang, Jingdong Chen |
ICASSP (2) | 3 |
| 2004 | An adaptive blind SIMO identification approach to joint multichannel time delay estimationabstractTime delay estimation (TDE) is a difficult problem in a reverberant environment and the traditional generalized cross-correlation (GCC) methods perform poorly. The adaptive eigenvalue decomposition (AED) algorithm recently proposed by the authors exploits a blind channel identification (BCI) technique and deals with room reverberation more effectively. The AED algorithm was developed for a two-channel system. It requires that the two channels do not share any common zeros (a necessary condition of system identifiability). We generalize the AED algorithm to multichannel (more than 2) systems. Compared to the AED algorithm, the generalized method is more robust since it is less likely for all channels to share a common zero when more sensors are used. Jingdong Chen, Yiteng Huang, Jacob Benesty |
ICASSP (4) | 1 |
| 2004 | Separating ISI and CCI in a two-step FIR Bezout equalizer for MIMO systems of frequency-selective channelsabstractThe use of multiple antennas at both the transmitter and the receiver in wireless communications implies a great channel capacity, but the signal detection is a crucial problem in achieving channel capacity, particularly in the general case of frequency selective channels, where both intersymbol interference (ISI) and cochannel interference (CCI) are significant. We show that ISI and CCI can be separated and then be cancelled in two different steps. We develop a two-step FIR Bezout equalizer and deduce the theoretically smallest length of its equalization filter. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (4) | 3 |
| 2004 | Recognition of noisy speech using dynamic spectral subband centroidsabstractDespite their widespread popularity as front-end parameters for speech recognition, the cepstral coefficients derived from either linear prediction analysis or a filter-bank are found to be sensitive to additive noise. In this letter, we discuss the use of spectral subband centroids for robust speech recognition. We show that centroids, if properly selected, can achieve recognition performance comparable to that of the mel-frequency cepstral coefficients (MFCCs) in clean speech, while delivering better performance than MFCC in noisy environments. A procedure is proposed to construct the dynamic centroid feature vector that essentially embodies the transitional spectral information. We discuss some properties of the proposed dynamic features. Jingdong Chen, Yiteng Huang, Kuldip K. Paliwal |
IEEE Signal Process. Lett. | 1 |
| 2004 | Time-delay estimation via linear interpolation and cross correlationabstractTime-delay estimation (TDE), which aims at measuring the relative time difference of arrival (TDOA) between different channels is a fundamental approach for identifying, localizing, and tracking radiating sources. Recently, there has been a growing interest in the use of TDE based locator for applications such as automatic camera steering in a room conferencing environment where microphone sensors receive not only the direct-path signal, but also attenuated and delayed replicas of the source signal due to reflections from boundaries and objects in the room. This multipath propagation effect introduces echoes and spectral distortions into the observation signal, termed as reverberation, which severely deteriorates a TDE algorithm in its performance. This paper deals with the TDE problem with emphasis on combating reverberation using multiple microphone sensors. The multichannel cross correlation coefficient (MCCC) is rederived here, in a new way, to connect it to the well-known linear interpolation technique. Some interesting properties and bounds of the MCCC are discussed and a recursive algorithm is introduced so that the MCCC can be estimated and updated efficiently when new data snapshots are available. We then apply the MCCC to the TDE problem. The resulting new algorithm can be treated as a natural generalization of the generalized cross correlation (GCC) TDE method to the multichannel case. It is shown that this new algorithm can take advantage of the redundancy provided by multiple microphone sensors to improve TDE against both reverberation and noise. Experiments confirm that the relative time-delay estimation accuracy increases with the number of sensors. Jacob Benesty, Jingdong Chen, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | A fast recursive algorithm for optimum sequential signal detection in a BLAST systemabstractBLAST (Bell Laboratories layered Space-Time) wireless systems are multiple-antenna communication schemes which can achieve very high spectral efficiencies in scattering environments, with no increase in bandwidth or transmitted power. The most popular and, by far, the most practical architecture is the so-called vertical BLAST (V-BLAST). The signal detection algorithm of a V-BLAST system is computationally very intensive. If the number of transmitters is M and is equal to the number of receivers, this complexity is proportional to M/sup 4/ at each sample time. In this paper, we propose a simple and very efficient algorithm that reduces the complexity by a factor of M. Jacob Benesty, Yiteng Huang, Jingdong Chen |
ICASSP (5) | 3 |
| 2003 | Robust time delay estimation exploiting spatial correlationabstractTo find the position of an acoustic source in a room, a set of relative delays among different microphone pairs has to be determined. The generalized cross-correlation method is the most popular to do so and is well explained in a landmark paper (Knapp and Carter (1996)). In this paper, we show how we can take advantage of the redundancy when more than two microphones are available. It is believed that the redundancy will help to better cope with noise and reverberation. The idea of cross-correlation coefficient between two signals is generalized to the multichannel case by using the notion of spatial prediction. The multichannel spatial correlation matrix is then deduced and it is shown how it can be used for time delay estimation. Jingdong Chen, Jacob Benesty, Yiteng Huang |
ICASSP (5) | 1 |
| 2003 | Adaptive blind identification of SIMO systems using channel cross-relation in the frequency domainabstractThe implementation of existing methods for blind identification of single-input multiple-output (SIMO) systems is limited in practice since they are difficult to execute in an adaptive mode and are, in general, computationally intensive. We extend our previous study (Huang, Y. and Benesty, J., Sig. Processing, vol.82, no.8, p.99-110, 2002) into the frequency domain and propose an unconstrained normalized multi-channel frequency-domain LMS (UNMCFLMS) algorithm. Numerical simulations show that the UNMCFLMS algorithm performs as well as (for a SIMO system with relatively short channel impulse responses) or better than (for a SIMO system with long channel impulse responses) its time-domain counterpart and the cross-relation (CR) batch method in practical situations. Yiteng Huang, Jacob Benesty, Jingdong Chen |
ICASSP (6) | 3 |
| 2003 | Cepstrum derived from differentiated power spectrum for robust speech recognition
Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
Speech Commun. | 1 |
| 2003 | Robust time delay estimation exploiting redundancy among multiple microphonesabstractTo find the position of an acoustic source in a room, typically, a set of relative delays among different microphone pairs needs to be determined. The generalized cross-correlation (GCC) method is the most popular to do so and is well explained in a landmark paper by Knapp and Carter. In this paper, the idea of cross-correlation coefficient between two random signals is generalized to the multichannel case by using the notion of spatial prediction. The multichannel spatial correlation matrix is then deduced and its properties are discussed. We then propose a new method based on the multichannel spatial correlation matrix for time delay estimation. It is shown that this new approach can take advantage of the redundancy when more than two microphones are available and this redundancy can help the estimator to better cope with noise and reverberation. Jingdong Chen, Jacob Benesty, Yiteng Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | Bell labs approach to Aurora evaluation on connected digit recognitionabstractABSTRACTIn this paper we study various front-endfeatures, mod-eling and adaptation algorithms on the Aurora 3 databases,including auditory, moment, and AM-FMmodulation fea-tures, context-dependentdigit models, segmental K-meanstraining, discriminative training, and model adaptations.The evaluation results on Aurora 3 are presented with abrief summary of our Aurora 2 results.1. INTRODUCTIONThe Aurora evaluation is for researchers to test their algo-rithms on noise robustness and compare results measuredon the same databases. So far, there are two tasks on theAurora evaluation, Aurora 2 and 3, both are for connecteddigit recognition. While the Aurora 2 databases use thecontrolled experiments by adding noise digitally to cleanEnglish digit strings [1], the Aurora 3 databases are col-lected in a real-worldcar environment in 4 languages. Inthis paper, we report our evaluation results on two of thelanguages, Spanish and German.2. BELL LABS APPROACHESIn this section, we present our baseline system then describethe different feature sets that have been used for this eval-uation. Alternative training strategies and acoustic modeladaptation techniques are also reviewed.A. Context-DependentModel: Similar to last year ap-proach [1], we have decided to use context-dependent(CD)digit models, together with Bell Labs recognition engine asbackend. This contrasts with the officialAurora backendthat is based on whole-worddigit models and the HTK en-gine. The officialbackend setup typically leads to poorerresults, especially in larger databases, and we believe that abetter baseline is beneficialto properly study the effect ofdifferent front-endson the finalrecognition performance.Last year, we investigated several approaches to buildCD digit models. Given the limited amount of trainingdata, especially in the Aurora3 databases, it is required torely on some tying techniques to build CD digit models.The Head-Body-Tail digit model structure (HBT) assumesthat CD digit models are built by concatenating a left-context-dependentunit (head) with a context-independentunit(body)followedbyaright-context-dependentunit(tail).For example, assuming that the lexicon contains 10 digitsplus a silence model, each digit model consists of a set of 1body, 11 heads and 11 tails (representing all left/right con-texts) [2]. We typically model each head and tail with a3-stateHMM, while a 4-stateHMM is used for each body.Most of the experiments done this year have been based onthe HBT structure. CD digit models can also be built as tri-phone modelsusing a decision tree. This is the approach weintroduced last year [1], and some of this year experimentshave been carried out using this model topology.B. Auditory Feature: The new auditory front-endinour recognition system was developed to mimic the robusthuman hearing in adverse acoustic environments [3, 4]. Inthe front-end,efficientsignal processing functions were im-plemented to satisfy both real-timeand computation costrequirements. Based on the analysis of the outer and mid-dle ear, a transfer function was constructed to replace thecommonly used preemphasis filter, and then a new set ofdigital auditory filters,which simulate auditory filteringinthe cochlea, replaces those used in the MFCC and PLP.The auditory feature extraction procedure consists of: anouter-middle-eartransfer function, FFT, frequency conver-sion from linear to the Bark scale, auditory filtering,non-linearity,and discrete cosine transform (DCT). In our previ-ous study[3], the feature has been evaluated in two tasks:connected-digit and large vocabulary, continuous speechrecognitionundervariousnoiseconditions,usingbothhand-setand hands-freedatainlandlineand wirelesstransmissionwith additive car and babble noise. Compared with theLPCC, MFCC, MEL-LPCC,and PLP features, the audi-tory feature achieved significantperformance improvement Jingdong Chen, Dimitris Dimitriadis, Hui Jiang 0001, Tor André Myrvoll, Olivier Siohan, Frank K. Soong |
INTERSPEECH | 1 |
| 2002 | Recognition of noisy speech using normalized moments
Jingdong Chen, Yiteng Huang, Frank K. Soong |
INTERSPEECH | 1 |
| 2001 | Sub-band based additive noise removal for robust speech recognitionabstractGriffith Sciences, Griffith School of Engineering Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2001 | Feature extraction and model-based noise compensation for noisy speech recognition evaluated on AURORA 2 taskabstractNo Full Text Kaisheng Yao, Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2000 | A block cosine transform and its application in speech recognition
Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 1999 | Regression class selection and speaker adaptation with MLLR in Mandarin continuous speech recognition
Chengrong Li, Jingdong Chen, Bo Xu 0002 |
EUROSPEECH | 2 |
| 1998 | A novel robust feature of speech signal based on the Mellin transform for speaker-independent speech recognitionabstractThis paper presents a novel kind of speech feature which is the modified Mellin transform of the log-spectrum of the speech signal (short for MMTLS). Because of the scale invariance property of the modified Mellin transform, the new feature is insensitive to the variation of the vocal tract length among individual speakers, and thus it is more appropriate for speaker-independent speech recognition than the popular used cepstrum. The preliminary experiments show that the performance of the MMTLS-based method is much better in comparison with those of the LPC- and MFC-based methods. Moreover, the error rate of this method is very consistent for different outlier speakers. Jingdong Chen, Bo Xu 0002, Taiyi Huang 0001 |
ICASSP | 1 |