Yiteng Huang

dblp:84/1272 · also Yiteng Arden Huang · DBLP profile ↗
← Back
70ranked-venue papers
24as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 19 first-author · 12 since 2021Artificial intelligence and machine learning · 23 · 5 first-author · 7 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MMW: Side Talk Rejection Multi-Microphone Whisper On Smart Glasses
abstract
Smart glasses are increasingly positioned as the nextgeneration interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in real-world noisy environments remains a major challenge, particularly due to interference from side speech. In this work, we introduce a novel side-talk rejection multimicrophone Whisper (MMW) framework for smart glasses, incorporating three key innovations. First, we propose a Mix Block based on a Tri-Mamba architecture to effectively fuse multi-channel audio at the raw waveform level, while maintaining compatibility with streaming processing. Second, we design a Frame Diarization Mamba Layer to enhance frame-level side-talk suppression, facilitating more efficient fine-tuning of Whisper models. Third, we employ a Multi-Scale Group Relative Policy Optimization (GRPO) strategy to jointly optimize frame-level and utterance-level side speech suppression. Experimental evaluations demonstrate that the proposed MMW system can reduce the word error rate (WER) by 4.95% in noisy conditions.
Yiteng Huang, Yangyang Shi, Saurabh Adya, Ming Sun 0013, Florian Metze
ASRU3
2025 Directional Source Separation for Robust Speech Recognition on Smart Glasses
abstract
Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We explore multiple beamformers to assist source separation by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. Furthermore, we investigate the training strategies for ASR when utilizing separated outputs. Our results suggest that jointly training a directional speech separation and ASR model achieves the best overall performance while balancing the wearer and conversation partner’s performance.
Tiantian Feng, Ju Lin, Yiteng Huang, Weipeng He, Kaustubh Kalgaonkar, Niko Moritz, Ming Sun 0013, Frank Seide
ICASSP3
2025 Effective Integration of KAN for Keyword Spotting
abstract
Keyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (KAN) can be used to enhance the performance of KWS. We explore various approaches to integrate KAN for a model architecture based on 1D Convolutional Neural Networks (CNN). We find that KAN is effective at modeling high-level features in lower-dimensional spaces, resulting in improved KWS performance when integrated appropriately. The findings shed light on understanding KAN for speech processing tasks and on other modalities for future researchers.
Anfeng Xu, Biqiao Zhang, Shuyu Kong, Yiteng Huang, Sangeeta Srivastava, Ming Sun 0013
ICASSP4
2025 M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
abstract
The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of unlabeled data. In this paper, we propose M-BEST-RQ, the first multi-channel speech foundation model for smart glasses, which is designed to leverage large-scale self-supervised learning (SSL) in an array-geometry agnostic approach. While prior work on multi-channel speech SSL only evaluated on simulated settings, we curate a suite of real downstream tasks to evaluate our model, namely (i) conversational automatic speech recognition (ASR), (ii) spherical active source localization, and (iii) glasses wearer voice activity detection, which are sourced from the MMCSG and EasyCom datasets. We show that a general-purpose M-BEST-RQ encoder is able to match or surpass supervised models across all tasks. For the conversational ASR task in particular, using only 8 hours of labeled speech, our model outperforms a supervised ASR baseline that is trained on 2000 hours of labeled data, which demonstrates the effectiveness of our approach.
Desh Raj, Ju Lin, Niko Moritz, Junteng Jia, Gil Keren, Egor Lakomkin, Yiteng Huang, Jacob Donley, Jay Mahadeokar, Ozlem Kalinli
ICASSP8
2025 Directional Speech Recognition with Full-Duplex Capability
Ju Lin, Yiteng Huang, Ming Sun 0013, Frank Seide, Florian Metze
INTERSPEECH2
2025 MASV: Speaker Verification with Global and Local Context Mamba
Yiteng Huang, Ming Sun 0013, Xinhao Mei, Yangyang Shi, Florian Metze
INTERSPEECH3
2025 Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
Jiamin Xie, Ju Lin, Yiteng Huang, Tyler Vuong, Zhaojiang Lin, Prashant Rawat, Sangeeta Srivastava, Ming Sun 0013, Florian Metze
INTERSPEECH3
2024 AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
abstract
Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise.When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses.This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.
Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie, Ming Sun 0013, Christian Fügen, Frank Seide
ICASSP3
2024 Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
Shuyu Kong, Biqiao Zhang, Yiteng Huang, Mumin Jin, Ming Sun 0013
INTERSPEECH5
2023 Disentangled Training with Adversarial Examples for Robust Small-Footprint Keyword Spotting
abstract
A keyword spotting (KWS) engine continuously running on the device is exposed to various speech signals that are usually unseen beforehand. It is a challenging problem to build a small-footprint and high-performing KWS model with robustness under different acoustic environments. In this paper, we explore how to effectively apply adversarial examples to improve KWS robustness. We propose datasource-aware disentangled learning with adversarial examples to reduce the mismatch between the original and adversarial data as well as the mismatch across original training datasources. The KWS model architecture is based on depth-wise separable convolution and a simple attention module. Experimental results demonstrate that the proposed learning strategy improves false reject rate by 40.31% at 1% false accept rate on the internal dataset, compared to the strongest baseline without adversarial examples. Our best-performing system achieves 98.06% accuracy on the Google Speech Commands V1 dataset.
Biqiao Zhang, Yiteng Huang, Shang-Wen Li 0001, Ming Sun 0013
ICASSP4
2023 Handling the Alignment for Wake Word Detection: A Comparison Between Alignment-Based, Alignment-Free and Hybrid Approaches
Vinicius Ribeiro, Yiteng Huang, Yuan Shangguan, Ming Sun 0013
INTERSPEECH2
2023 A Novel Traffic Classifier With Attention Mechanism for Industrial Internet of Things
abstract
With the development of the Industrial Internet of Things (IIoT), the complex traffic generated by large-scale IIoT devices presents challenges for traffic analysis. Most of existing deep learning-based traffic analysis methods use a single flow for classification, resulting in being misled by the irrelevant flow. Thus, it is necessary to use flow sequences for traffic analysis. However, existing models fail to effectively distinguish unimportant flows in flow sequence, which affects the classification performance. To address the aforementioned challenges, we propose a novel traffic classifier called flow transformer to perform traffic analysis with flow sequences, which leverages multihead attention mechanism to strengthen the information interaction between related flows. Besides, the RF-based feature selection method is designed to select the optimal feature combination, avoiding insignificant features from reducing the performance of the classifier. Experimental results on three real-world traffic datasets demonstrate that our method outperforms state-of-the-art methods with a large margin.
Ruijie Zhao 0001, Yiteng Huang, Xianwen Deng, Yong Shi 0009, Jiabin Li, Zijing Huang, Zhi Xue
IEEE Trans. Ind. Informatics2
2021 Personalized Keyphrase Detection Using Speaker and Environment Information
abstract
In this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The system is implemented with an end-to-end trained automatic speech recognition (ASR) model and a text-independent speaker verification model. To address the challenge of detecting these keyphrases under various noisy conditions, a speaker separation model is added to the feature frontend of the speaker verification model, and an adaptive noise cancellation (ANC) algorithm is included to exploit cross-microphone noise coherence. Our experiments show that the text-independent speaker verification model largely reduces the false triggering rate of the keyphrase detection, while the speaker separation model and adaptive noise cancellation largely reduce false rejections.
Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Ding Zhao, Yiteng Huang, Arun Narayanan, Ian McGraw
Interspeech6
2021 Flow Transformer: A Novel Anonymity Network Traffic Classifier with Attention Mechanism
abstract
Supervising anonymity network is a critical issue in the field of network security, and traditional traffic analysis methods cannot cope with complex anonymity traffic. In recent years, the traffic analysis method based on deep learning has achieved good performance. However, most of the existing studies do not consider the temporal-spatial correlation of the traffic, and only use a single flow for classification. A few works take continuous flows as flow sequence for traffic classification, but they do not distinguish the different importance of each flow. To tackle this issue, we propose a novel flow-based traffic classifier called FLOW TRANSFORMER to classify anonymity network traffic. FLOW TRANSFORMER uses multi-head attention mechanism to set higher weights for important flows, and extracts flow sequence features according to the importance weights. Besides, the RF-based feature selection method is designed to select the optimal feature combination, which can effectively avoid the insignificant features from reducing the performance and efficiency of the classifier. Experimental results on two real-world traffic datasets demonstrate that the proposed method outperforms state-of-the-art methods with a large margin.
Ruijie Zhao 0001, Yiteng Huang, Xianwen Deng, Zhi Xue, Jiabin Li, Zijing Huang
MSN2
2019 Hotword Cleaner: Dual-microphone Adaptive Noise Cancellation with Deferred Filter Coefficients for Robust Keyword Spotting
abstract
This paper presents a novel dual-microphone speech enhancement algorithm to improve noise robustness of hotword (wake-word) detection as a special application of keyword spotting. It exploits two unique properties of hotwords: they are leading phrases of valid voice queries that we intend to respond and have short durations. Consequently an STFT-based adaptive noise cancellation method modified to use deferred filter coefficients is proposed to extract hotwords out from noisy stereo microphone signals. The new algorithm is tested with two considerably different neural hotword detectors. Both systems have significantly reduced the false-reject rate when background has strong TV noise.
Yiteng Huang, Turaj Zakizadeh Shabestary, Alexander Gruenstein
ICASSP1
2019 Multi-Microphone Adaptive Noise Cancellation for Robust Hotword Detection
Yiteng Huang, Turaj Zakizadeh Shabestary, Alexander Gruenstein
INTERSPEECH1
2019 Optimal Power Allocation for SCMA Downlink Systems Based on Maximum Capacity
abstract
Sparse code multiple access (SCMA) is a novel type of non-orthogonal multiple access technology that combines the concepts of CDMA and OFDMA. The advantages of SCMA include high capacity, low time delay, and high date rate. In this paper, a power allocation algorithm is proposed for SCMA downlink systems where each tone is taken by more than one user to maximize the system's sum capacity. In SCMA systems, users are divided into different user groups. Thus, our proposed algorithm includes three-level power allocation. Since the power allocation problem is non-convex, the complexity of finding the optimal solutions is prohibitive. The Lagrange dual decomposition method is employed to efficiently solve the non-convex optimization problem. Results show that the optimized algorithm can significantly improve the sum capacity.
Shuai Han 0002, Yiteng Huang, Weixiao Meng 0001, Cheng Li 0005, Dageng Chen
IEEE Trans. Commun.2
2018 Supervised Noise Reduction for Multichannel Keyword Spotting
abstract
This paper presents a robust, small-footprint, far-field keyword spotting (KWS) algorithm, which was inspired by the human auditory system's ability to achieve the so-called cocktail party effect in adverse acoustic environments. It introduces the idea of combining microphone-array speech enhancement with machine learning, by incorporating a feedback path from the neural network (NN) KWS classifier to its signal preprocessing frontend so that frontend noise reduction can benefit from, and in turn, better serve backend machine intelligence. We find that the new system can significantly improve KWS performance for Google Home when there is strong music or TV noise in the background. While this innovative and successfully validated strategy of combining signal processing and machine learning is developed for KWS, its technical feasibility is presumably extensible to many other applications, including noise robust speaker identification and automatic speech recognition.
Yiteng Huang, Thad Hughes, Turaj Zakizadeh Shabestary, Taylor Applebaum
ICASSP1
2018 Power Allocation for SCMA Downlink Systems Based on Maximum Energy Efficiency
abstract
Sparse code multiple access(SCMA) is a novel kind of non-orthogonal multiple access technology which has the advantages of supporting more connections, low time delay and overcomes the near-far effect in CDMA system. In this paper we propose power allocation algorithms for SCMA downlink systems to maximize the energy efficiency. Since the resource allocation problem is non-convex, we employ the Lagrange dual decomposition method and Dinkelbach theory to solve the optimization problem. Our results shows that energy efficiency can be improved significantly by adopting optimal power allocation method. The SCMA maximum energy efficiency system save more energy in the condition of satisfying the user QoS requirement.
Yiteng Huang, Shuai Han 0002, Shizeng Guo, Weixiao Meng 0001, Cheng Li 0005
IWCMC1
2017 Practically efficient nonlinear acoustic echo cancellers using cascaded block RLS and FLMS adaptive filters
abstract
This paper presents a practically efficient implementation for non-linear acoustic echo cancellation (NAEC). The echo path is modeled by a novel hybrid Taylor-Volterra pre-processor followed by a linear FIR filter. A cascaded block RLS and unconstrained FLMS adaptive algorithm is developed to jointly identify the pre-processor and the FIR filter. This implementation is validated via simulations.
Yiteng Huang, Jan Skoglund, Alejandro Luebs
ICASSP1
2016 Globally optimized least-squares post-filtering for microphone array speech enhancement
abstract
Existing post-filtering techniques for microphone array speech enhancement have two common deficiencies. First, they assume that the noise is either white or diffuse and cannot deal with point inter-ferers. Second, they estimate the post-filter coefficients using only two microphones at a time and then perform averaging over all microphone pairs, yielding a suboptimal solution at best. In this paper, we present a novel post-filtering algorithm that alleviates the first limitation by using a more generalized signal model including not only white and diffuse but also point interferers, and overcomes the second deficiency by offering a globally optimized least-squares solution over all microphones. It is shown by simulations that the proposed method outperforms the existing algorithms in many different acoustic scenarios.
Yiteng Huang, Alejandro Luebs, Jan Skoglund, W. Bastiaan Kleijn
ICASSP1
2012 A Perspective on Differential Microphone Arrays in the Context of Noise Reduction
abstract
In this correspondence, we study the performance of differential microphone arrays (DMAs) in terms of noise reduction, speech distortion, and signal-to-noise ratio (SNR) gain. We also investigate their beampatterns and array gains. We start by establishing the expressions of these performance measures involving general derivatives of the channel transfer functions. Afterwards, we specify our results in the case of anechoic near-field and far-field propagation models.
Jacob Benesty, Mehrez Souden, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2012 A Multi-Frame Approach to the Frequency-Domain Single-Channel Noise Reduction Problem
abstract
This paper focuses on the class of single-channel noise reduction methods that are performed in the frequency domain via the short-time Fourier transform (STFT). The simplicity and relative effectiveness of this class of approaches make them the dominant choice in practical systems. Over the past years, many popular algorithms have been proposed. These algorithms, no matter how they are developed, have one feature in common: the solution is eventually formulated as a gain function applied to the STFT of the noisy signal only in the current frame, implying that the interframe correlation is ignored. This assumption is not accurate for speech enhancement since speech is a highly self-correlated signal. In this paper, by taking the interframe correlation into account, a new linear model for speech spectral estimation and some optimal filters are proposed. They include the multi-frame Wiener and minimum variance distortionless response (MVDR) filters. With these filters, both the narrowband and fullband signal-to-noise ratios (SNRs) can be improved. Furthermore, with the MVDR filter, speech distortion at the output can be zero. Simulations present promising results in support of the claimed merits obtained by theoretical analysis.
Yiteng Huang, Jacob Benesty
IEEE Trans. Speech Audio Process.1
2011 A single-channel noise reduction MVDR filter
abstract
Most existing approaches for single-channel noise reduction in the frequency domain via the short-time Fourier transform (STFT) assume that consecutive time-frames are uncorrelated with each other. As a result, algorithms based on this assumption do not, obviously, take the interframe correlation into account. In this paper, we pro pose a framework that considers this interframe correlation. An important consequence in including this useful information, is that now it is possible to derive a (single-channel) minimum variance distortionless response (MVDR) filter for noise reduction. The experimental study shows impressive results.
Jacob Benesty, Yiteng Huang
ICASSP2
2011 On single-channel noise reduction in the time domain
abstract
In this paper, we revisit the noise-reduction problem in the time domain and present a way to decompose the filtered speech into two uncorrelated (orthogonal) components: the desired speech and the interference. Based on this new decomposition, we discuss how to form different optimization cost functions and address the issue of how to design different noise-reduction filters by optimizing these new cost functions. Particularly, we cover the design of the maximum signal-to-noise-ratio (SNR), the Wiener, the minimum variance distortionless response (MVDR), and the tradeoff filters. It is interesting that with this new decomposition, we can now design the MVDR filter that can achieve noise reduction without adding speech distortion in the single-channel case, which has never been seen before. We also demonstrate that the maximum SNR, Wiener, and tradeoff filters are identical to the MVDR filter up to a scaling factor. From a theoretical point of view, this scaling factor is not significant and should not affect the output SNR at any processing time. But from a practical viewpoint, the scaling factor can be time-varying due to the nonstationarity of the speech and possibly the noise and can cause discontinuity in the residual noise level, which is unpleasant to listen to. As a result, it is essential to have the scaling factor right from one processing sample (or frame) to another in order to avoid large distortions and for this reason, it is recommended to use the MVDR filter in speech enhancement applications.
Jingdong Chen, Jacob Benesty, Yiteng Huang, Tomas Gänsler
ICASSP3
2011 Binaural Noise Reduction in the Time Domain With a Stereo Setup
abstract
Binaural noise reduction with a stereophonic (or simply stereo) setup has become a very important problem as stereo sound systems and devices are being more and more deployed in modern voice communications. This problem is very challenging since it requires not only the reduction of the noise at the stereo inputs, but also the preservation of the spatial information embodied in the two channels so that after noise reduction the listener can still localize the sound source from the binaural outputs. As a result, simply applying a traditional single-channel noise reduction technique to each channel individually may not work as the spatial effects may be destroyed. In this paper, we present a new formulation of the binaural noise reduction problem in stereo systems. We first form a complex signal from the stereo inputs with one channel being its real part and the other being its imaginary part. By doing so, the binaural noise reduction problem can be processed by a single-channel widely linear filter. The widely linear estimation theory is then used to derive optimal noise reduction filters that can fully take advantage of the noncircularity of the complex speech signal to achieve noise reduction while preserving the desired signal (speech) and spatial information. With this new formulation, the Wiener, minimum variance distortionless response (MVDR), maximum signal-to-noise ratio (SNR), and tradeoff filters are derived. Experiments are provided to justify the effectiveness of these filters.
Jacob Benesty, Jingdong Chen, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2010 Study of the widely linear Wiener filter for noise reduction
abstract
This paper develops a new widely linear noise-reduction Wiener filter based on the variance and pseudo-variance of the short-time Fourier transform coefficients of speech signals. We show that this new noise-reduction filter has many interesting properties, including but not limited to: 1) it causes less speech distortion as compared to the classical noise-reduction Wiener filter; 2) its minimum mean-squared error (MSE) is smaller than that of the classical Wiener filter; 3) it can increase the subband signal-to-noise ratio (SNR), while the classical Wiener filter has no effect on the subband SNR for any given signal frame and subband.
Jacob Benesty, Jingdong Chen, Yiteng Huang
ICASSP3
2010 Analysis of the frequency-domain Wiener filter with the prediction gain
abstract
This paper presents a theoretical analysis on the performance of the optimal noise-reduction filter in the frequency domain. Using the autoregressive (AR) model to model both the clean speech and noise, we build the relationship between the Wiener filter and the AR parameters of the clean speech and noise signals. We show that if noise is not predictable, the Wiener filter is mostly related to the AR parameters of the desired speech signal. On the contrary, if the desired signal is not predictable, the Wiener filter is then mostly related to the AR parameters of the noise signal. More importantly, we provide the bounds for noise reduction, speech distortion, and SNR improvement, and show that the performance of the Wiener filter in terms of SNR improvement and degree of noise reduction and speech distortion is closely related to the prediction gain of the desired speech and noise signals.
Jingdong Chen, Jacob Benesty, Yiteng Huang
ICASSP3
2010 On widely linear Wiener and tradeoff filters for noise reduction
Jacob Benesty, Jingdong Chen, Yiteng Huang
Speech Commun.3
2010 A Widely Linear Distortionless Filter for Single-Channel Noise Reduction
abstract
Traditionally in the single-channel noise-reduction problem, speech distortion is inevitable since the desired signal is also filtered while filtering the noise. In fact, the more the noise is reduced, the more the speech distortion is added into the desired signal, as proved in the literature. So, if we require no speech distortion, we either end up with no noise reduction at all or have to use multiple sensors. In this paper, we attempt to apply the widely linear (WL) estimation theory to noise reduction. Unlike the traditional approaches that only filter the short-time Fourier transform (STFT) of the noisy signal, the method developed in this paper applies the noise-reduction filter to both the STFT of the noisy signal and its conjugate. With the constraint of no speech distortion, a WL distortionless filter is derived. We show that this new optimal filter can fully take advantage of the noncircularity property of speech signals to achieve up to 3-dB signal-to-noise-ratio (SNR) improvement without introducing any speech distortion, which can only be obtained with the traditional approaches if two or more microphones are used.
Jacob Benesty, Jingdong Chen, Yiteng Huang
IEEE Signal Process. Lett.3
2009 On noise reduction in the Karhunen-Loève expansion domain
abstract
In this paper, we study the noise-reduction problem in the Karhunen-Loève expansion domain. We develop two classes of optimal filters. The first class estimates a frame of speech by filtering the corresponding frame of the noisy speech. We will show that several well-known existing methods belong or are closely related to this category. The second class, which has not been studied before, obtains noise reduction by filtering not only the current frame, but also a number of previous consecutive frames of the noisy speech. We will discuss how to design the optimal noise-reduction filters in each class and demonstrate the properties of the deduced optimal filters.
Jacob Benesty, Jingdong Chen, Yiteng Huang
ICASSP3
2009 Using the Pearson correlation coefficient to develop an optimally weighted cross relation based blind SIMO identification algorithm
abstract
Blind SIMO identification is challenging when additive noise is strong and for ill-conditioned/acoustic SIMO systems. A weighted cross relation (CR) algorithm presumably can be robust to noise but there lacks a practical way to define the weights. In this paper, the Pearson correlation coefficient (PCC) is used to develop an optimally weighted CR algorithm, which is validated by simulations.
Yiteng Huang, Jacob Benesty, Jingdong Chen
ICASSP1
2009 Noise Reduction Algorithms in a Generalized Transform Domain
abstract
Noise reduction for speech applications is often formulated as a digital filtering problem, where the clean speech estimate is obtained by passing the noisy speech through a linear filter/transform. With such a formulation, the core issue of noise reduction becomes how to design an optimal filter (based on the statistics of the speech and noise signals) that can significantly suppress noise without introducing perceptually noticeable speech distortion. The optimal filters can be designed either in the time or in a transform domain. The advantage of working in a transform space is that, if the transform is selected properly, the speech and noise signals may be better separated in that space, thereby enabling better filter estimation and noise reduction performance. Although many different transforms exist, most efforts in the field of noise reduction have been focused only on the Fourier and Karhunen-Loeve transforms. Even with these two, no formal study has been carried out to investigate which transform can outperform the other. In this paper, we reformulate the noise reduction problem into a more generalized transform domain. We will show some of the advantages of working in this generalized domain, such as 1) different transforms can be used to replace each other without any requirement to change the algorithm (optimal filter) formulation, and 2) it is easier to fairly compare different transforms for their noise reduction performance. We will also address how to design different optimal and suboptimal filters in such a generalized transform domain.
Jacob Benesty, Jingdong Chen, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2009 Study of the Noise-Reduction Problem in the Karhunen-LoÈve Expansion Domain
abstract
Noise reduction, which aims at estimating a clean speech from a noisy observation, has long been an active research area. The standard approach to this problem is to obtain the clean speech estimate by linearly filtering the noisy signal. The core issue, then, becomes how to design an optimal linear filter that can significantly suppress noise without introducing perceptually noticeable speech distortion. Traditionally, the optimal noise-reduction filters are formulated in either the time or the frequency domains. This paper studies the problem in the Karhunen–LoÈve expansion domain. We develop two classes of optimal filters. The first class achieves a frame of speech estimate by filtering the corresponding frame of the noisy speech. We will show that many existing methods such as the widely used Wiener filter and subspace technique are closely related to this category. The second class obtains noise reduction by filtering not only the current frame, but also a number of previous consecutive frames of the noisy speech. We will discuss how to design the optimal noise-reduction filters in each class and demonstrate, through both theoretical analysis and experiments, the properties of the deduced optimal filters.
Jingdong Chen, Jacob Benesty, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2008 A minimum speech distortion multichannel algorithm for noise reduction
abstract
Noise reduction using multiple microphones remains a challenging and crucial research problem. This paper presents a new multichannel noise-reduction algorithm based on spatio-temporal prediction. Unlike many multichannel techniques that attempt to achieve both speech dereverberation and noise reduction at the same time, this new approach puts aside speech dereverberation and formulates the problem as one of estimating the speech component received at one microphone using the observations from all the available microphones. In comparison with the existing techniques such as beamforming, this new multichannel approach has many appealing properties: it does not require the knowledge of the source location or the channel impulse responses; the multiple microphones do not have to be arranged into a specific array geometry; it works the same for both the far-field and near-field cases; and most importantly, it can produce very good noise reduction with minimum speech distortion in real acoustic environments.
Jacob Benesty, Jingdong Chen, Yiteng Huang
ICASSP3
2008 Generalized crosstalk cancellation and equalization using multiple loudspeakers for 3D sound reproduction at the ears of multiple listeners
abstract
A classical crosstalk cancellation and equalization (CTCE) system uses two loudspeakers and assumes only one listener. In this paper, the idea is generalized to the design of using multiple loudspeakers for multiple listeners. We will show that if the number of loudspeakers is equal to the number of ears, only a least-squares (LS) solution can be obtained, while using more loudspeakers than ears we have more options: either an LS solution or an exact solution for perfect CTCE. Via simulations using real impulse responses measured in the varechoic chamber at Bell Labs, we learn that a CTCE system employing more loudspeakers will be more robust to errors in the estimated acoustic impulse responses.
Yiteng Huang, Jacob Benesty, Jingdong Clien
ICASSP1
2008 On the Importance of the Pearson Correlation Coefficient in Noise Reduction
abstract
Noise reduction, which aims at estimating a clean speech from noisy observations, has attracted a considerable amount of research and engineering attention over the past few decades. In the single-channel scenario, an estimate of the clean speech can be obtained by passing the noisy signal picked up by the microphone through a linear filter/transformation. The core issue, then, is how to find an optimal filter/transformation such that, after the filtering process, the signal-to-noise ratio (SNR) is improved but the desired speech signal is not noticeably distorted. Most of the existing optimal filters (such as the Wiener filter and subspace transformation) are formulated from the mean-square error (MSE) criterion. However, with the MSE formulation, many desired properties of the optimal noise-reduction filters such as the SNR behavior cannot be seen. In this paper, we present a new criterion based on the Pearson correlation coefficient (PCC). We show that in the context of noise reduction the squared PCC (SPCC) has many appealing properties and can be used as an optimization cost function to derive many optimal and suboptimal noise-reduction filters. The clear advantage of using the SPCC over the MSE is that the noise-reduction performance (in terms of the SNR improvement and speech distortion) of the resulting optimal filters can be easily analyzed. This shows that, as far as noise reduction is concerned, the SPCC-based cost function serves as a more natural criterion to optimize as compared to the MSE.
Jacob Benesty, Jingdong Chen, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2008 A Minimum Distortion Noise Reduction Algorithm With Multiple Microphones
abstract
The problem of noise reduction using multiple microphones has long been an active area of research. Over the past few decades, most efforts have been devoted to beamforming techniques, which aim at recovering the desired source signal from the outputs of an array of microphones. In order to work reasonably well in reverberant environments, this approach often requires such knowledge as the direction of arrival (DOA) or even the room impulse responses, which are difficult to acquire reliably in practice. In addition, beamforming has to compromise its noise reduction performance in order to achieve speech dereverberation at the same time. This paper presents a new multichannel algorithm for noise reduction, which formulates the problem as one of estimating the speech component observed at one microphone using the observations from all the available microphones. This new approach explicitly uses the idea of spatial–temporal prediction and achieves noise reduction in two steps. The first step is to determine a set of inter-sensor optimal spatial–temporal prediction transformations. These transformations are then exploited in the second step to form an optimal noise-reduction filter. In comparison with traditional beamforming techniques, this new method has many appealing properties: it does not require DOA information or any knowledge of either the reverberation condition or the channel impulse responses; the multiple microphones do not have to be arranged into a specific array geometry; it works the same for both the far-field and near-field cases; and, most importantly, it can produce very good and robust noise reduction with minimum speech distortion in practical environments. Furthermore, with this new approach, it is possible to apply postprocessing filtering for additional noise reduction when a specified level of speech distortion is allowed.
Jingdong Chen, Jacob Benesty, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2008 Analysis and Comparison of Multichannel Noise Reduction Methods in a Common Framework
abstract
Noise reduction for speech enhancement is a useful technique, but in general it is a challenging problem. While a single-channel algorithm is easy to use in practice, it inevitably introduces speech distortion to the desired speech signal while reducing noise. Today, the explosive growth in computational power and the continuous drop in the cost and size of acoustic electric transducers are driving the interest of employing multiple microphones in speech processing systems. This opens new opportunities for noise reduction. In this paper, we present an analysis of three multichannel noise reduction algorithms, namely Wiener filter, subspace, and spatial-temporal prediction, in a common framework. We intend to investigate whether it is possible for the multichannel noise reduction algorithms to reduce noise without speech distortion. Finally, we justify what we learn via theoretical analyses by simulations using real impulse responses measured in the varechoic chamber at Bell Labs.
Yiteng Huang, Jacob Benesty, Jingdong Chen
IEEE Trans. Speech Audio Process.1
2007 On Recursive and Fast Recursive Computation of the Capon Spectrum
abstract
The Capon spectrum, which is known to have better resolution than the periodogram, has been widely used in various applications. Normally, the Capon spectrum is estimated through the direct computation of the inverse of the data correlation (or covariance) matrix. This so-called direct inverse approach is, however, computationally very expensive due to the high computational cost involved in the matrix inversion. This paper deals with fast and efficient algorithms in computing the Capon spectrum. Inspired from the recursive idea established in the area of adaptive signal processing, we first derive a recursive Capon algorithm. This new algorithm does not require an explicit matrix inversion, and hence is more efficient to implement than the direct inverse method. We then develop a fast version of the recursive algorithm, which can further reduce the complexity of the recursive one by an order of magnitude.
Jacob Benesty, Jingdong Chen, Yiteng Huang
ICASSP (3)3
2007 An Acoustic MIMO Framework for Analyzing Microphone-Array Beamforming
abstract
Although a significant amount of research attention has been devoted to microphone-array beamforming, the performance of all the developed algorithms in practical acoustic environments is still far from meeting our expectation. So further research efforts on this topic are indispensable. In this paper, we treat a microphone array as a multiple-input multiple-output (MIMO) system and develop a general framework for analyzing performance of beamforming algorithms based on the acoustic MIMO channel impulse responses. Under this framework, we study the bounds for the length of beamforming filter, which in turn shows the performance bounds of beamforming in terms of speech dereverberation and interference suppression. We also discuss the intrinsic relationships among different classical beamforming techniques and explain, from the channel condition point of view, what the prerequisites have to be fulfilled in order for those techniques to work.
Jingdong Chen, Jacob Benesty, Yiteng Huang
ICASSP (1)3
2007 Laplace Entropy and its Application to Time Delay Estimation for Speech Signals
abstract
Time delay estimation (TDE) is a basic technique for numerous applications where there is a need to localize and track a radiating source. It is particularly challenging in the presence of noise and reverberation, and when the source signal is speech which is inherently nonstationary and random. The most important TDE algorithms for two sensors are based on the generalized cross-correlation (GCC) method. These algorithms perform reasonably well when reverberation or noise is not too high. In an earlier study of the authors, a more sophisticated approach was proposed. It employs more sensors and takes advantage of their delay redundancy to improve the precision of the TDOA (time difference of arrival) estimate between the first two sensors. The approach is based on the multichannel cross-correlation coefficient (MCCC) and was found more robust to noise and reverberation. In this paper, we show that this approach can also be developed on a basis of joint entropy. For Gaussian signals, we show that, in the search of the TDOA estimate, maximizing MCCC is equivalent to minimizing joint entropy. But with the generalization of the idea to non-Gaussian speech signals, the joint entropy based new multichannel TDE algorithm manifests a potential to outperform the MCCC-based method. Since there is no rigorous mathematical formula for speech entropy, we use the assumption that speech can be plausibly modeled by a Laplace distribution and develop a practical approximation of Laplace entropy for TDE of speech signals. The performance of the proposed new algorithm is investigated via simulations.
Yiteng Huang, Jacob Benesty, Jingdong Chen
ICASSP (1)1
2007 On the optimal linear filtering techniques for noise reduction
Jingdong Chen, Jacob Benesty, Yiteng Huang
Speech Commun.3
2007 Time Delay Estimation via Minimum Entropy
abstract
Time delay estimation (TDE) is a basic technique for numerous applications where there is a need to localize and track a radiating source. The most important TDE algorithms for two sensors are based on the generalized cross-correlation (GCC) method. These algorithms perform reasonably well when reverberation or noise is not too high. In an earlier study by the authors, a more sophisticated approach was proposed. It employs more sensors and takes advantage of their delay redundancy to improve the precision of the time difference of arrival (TDOA) estimate between the first two sensors. The approach is based on the multichannel cross-correlation coefficient (MCCC) and was found more robust to noise and reverberation. In this letter, we show that this approach can also be developed on a basis of joint entropy. For Gaussian signals, we show that, in the search of the TDOA estimate, maximizing MCCC is equivalent to minimizing joint entropy. However, with the generalization of the idea to non-Gaussian signals (e.g., speech), the joint entropy-based new TDE algorithm manifests a potential to outperform the MCCC-based method
Jacob Benesty, Yiteng Huang, Jingdong Chen
IEEE Signal Process. Lett.2
2007 On Crosstalk Cancellation and Equalization With Multiple Loudspeakers for 3-D Sound Reproduction
abstract
People prefer to be able to enjoy spatial audio without wearing a headphone. Such a tethered device is anyway inconvenient and undesirable, if not cumbersome. Alternatively, 3D sound can be delivered to a listener with loudspeakers. However, crosstalk arises, and the rendered binaural signals are distorted by room reverberation when arriving at the listener's two ears, which lead to the need for a crosstalk cancellation and equalization (CTCE) system. Classical CTCE systems employ only two loudspeakers, and their performance is usually unsatisfactory in practice. While the idea of using more loudspeakers has been investigated, it was never shown why using more loudspeakers is theoretically more advantageous for CTCE. In this letter, we will study this problem and demonstrate that with two loudspeakers, only a least-squares (LS) solution can be obtained, while using multiple loudspeakers, we have more options: either an LS solution or an exact solution for perfect CTCE. These findings are justified by simulations using real impulse responses measured in the varechoic chamber at Bell Labs.
Yiteng Huang, Jacob Benesty, Jingdong Chen
IEEE Signal Process. Lett.1
2007 On Microphone-Array Beamforming From a MIMO Acoustic Signal Processing Perspective
abstract
Although many microphone-array beamforming algorithms have been developed over the past few decades, most such algorithms so far can only offer limited performance in practical acoustic environments. The reason behind this has not been fully understood and further research on this matter is indispensable. In this paper, we treat a microphone array as a multiple-input multiple-output (MIMO) system and study its signal-enhancement performance. Our major contribution is fourfold. First, we develop a general framework for analyzing performance of beamforming algorithms based on the acoustic MIMO channel impulse responses. Second, we study the bounds for the length of the beamforming filter, which in turn shows the performance bounds of beamforming in terms of speech dereverberation and interference suppression. Third, we address the connection between beamforming and the multiple-input/output inverse theorem (MINT). Finally, we discuss the intrinsic relationships among different classical beamforming techniques and explain, from the channel condition perspective, what the prerequisites are for those techniques to work.
Jacob Benesty, Jingdong Chen, Yiteng Huang, Jacek Dmochowski
IEEE Trans. Speech Audio Process.3
2006 Estimation of the Coherence Function with the MVDR Approach
abstract
The minimum variance distortionless response (MVDR), originally developed by Capon for frequency-wavenumber analysis, is a very well established method in array processing. It is also used in spectral estimation. The aim of this paper is to show how the MVDR method can be used to estimate the magnitude squared coherence (MSC) function, which is very useful in so many applications but so few methods exist to estimate it. Simulations show that our algorithm gives much more reliable results than the one based on the popular Welch's method.
Jacob Benesty, Jingdong Chen, Yiteng Huang
ICASSP (3)3
2006 Speech Acquisition and Enhancement in a Reverberant, Cocktail-Party-Like Environment
abstract
Developing a successful multi-microphone speech acquisition system in a reverberant, cocktail-party-like environment is a very challenging problem since both interfering sources and reverberation need to be well controlled. In this paper, we propose an algorithm based on blind SIMO identification. We first blindly identify the channels from the interfering sources to all the microphones. Then we extract the speech signal of interest. Finally speech dereverberation is performed using the MINT method. Simulations with acoustic impulse responses measured in the varechoic chamber at Bell Labs are carried out to verify the proposed algorithm.
Yiteng Huang, Jacob Benesty, Jingdong Chen
ICASSP (5)1
2006 Identification of acoustic MIMO systems: Challenges and opportunities
Yiteng Huang, Jacob Benesty, Jingdong Chen
Signal Process.1
2006 New insights into the noise reduction Wiener filter
abstract
The problem of noise reduction has attracted a considerable amount of research attention over the past several decades. Among the numerous techniques that were developed, the optimal Wiener filter can be considered as one of the most fundamental noise reduction approaches, which has been delineated in different forms and adopted in various applications. Although it is not a secret that the Wiener filter may cause some detrimental effects to the speech signal (appreciable or even significant degradation in quality or intelligibility), few efforts have been reported to show the inherent relationship between noise reduction and speech distortion. By defining a speech-distortion index to measure the degree to which the speech signal is deformed and two noise-reduction factors to quantify the amount of noise being attenuated, this paper studies the quantitative performance behavior of the Wiener filter in the context of noise reduction. We show that in the single-channel case the a posteriori signal-to-noise ratio (SNR) (defined after the Wiener filter) is greater than or equal to the a priori SNR (defined before the Wiener filter), indicating that the Wiener filter is always able to achieve noise reduction. However, the amount of noise reduction is in general proportional to the amount of speech degradation. This may seem discouraging as we always expect an algorithm to have maximal noise reduction without much speech distortion. Fortunately, we show that speech distortion can be better managed in three different ways. If we have some a priori knowledge (such as the linear prediction coefficients) of the clean speech signal, this a priori knowledge can be exploited to achieve noise reduction while maintaining a low level of speech distortion. When no a priori knowledge is available, we can still achieve a better control of noise reduction and speech distortion by properly manipulating the Wiener filter, resulting in a suboptimal Wiener filter. In case that we have multiple microphone sensors, the multiple observations of the speech signal can be used to reduce noise with less or even no speech distortion.
Jingdong Chen, Jacob Benesty, Yiteng Huang, Simon Doclo
IEEE Trans. Speech Audio Process.3
2005 Time delay estimation via multichannel cross-correlation [audio signal processing applications]
abstract
Time delay estimation (TDE) in a reverberant acoustical environment is a very challenging and difficult problem. This paper tackles the problem by exploiting the redundant information provided by multiple microphone sensors. To do so, the multichannel crosscorrelation coefficient (MCCC) is re-derived, in a new way, to connect it to the well-known linear interpolation technique. Some interesting properties and bounds of MCCC are discussed, and a recursive algorithm is then introduced so that MCCC can be estimated and updated efficiently when new data snapshots are available. We then apply the MCCC to the TDE problem, resulting a multichannel cross-correlation algorithm that can be treated as a natural generalization of the generalized cross-correlation (GCC) TDE method to the multichannel case. It is shown that this method can take advantage of the redundancy provided by multiple microphone sensors to improve TDE against both reverberation and noise.
Jingdong Chen, Yiteng Huang, Jacob Benesty
ICASSP (3)2
2005 Adaptive blind SIMO identification: derivation of an optimal step size for the unconstrained multichannel LMS algorithm
abstract
Adaptive algorithms for blindly identifying SIMO systems are appealing because of their computational efficiency and capability of continuously tracking a time-varying system. Adaptive multichannel LMS (MCLMS) algorithms (with and without the unit-norm constraint) are analyzed and the optimal step size is derived. A simple yet effective variable step-size unconstrained MCLMS algorithm is proposed and its performance is evaluated with simulations.
Yiteng Huang, Jacob Benesty, Jingdong Chen
ICASSP (3)1
2005 A generalized MVDR spectrum
abstract
The minimum variance distortionless response (MVDR) approach is very popular in array processing. It is also employed in spectral estimation where the Fourier matrix is used in the optimization process. First, we give a general form of the MVDR where any unitary matrix can be used to estimate the spectrum. Second and most importantly, we show how the MVDR method can be used to estimate the magnitude squared coherence function, which is very useful in so many applications but so few methods exist to estimate it. Simulations show that our algorithm gives much more reliable results than the one based on the popular Welch's method.
Jacob Benesty, Jingdong Chen, Yiteng Huang
IEEE Signal Process. Lett.3
2005 Optimal step size of the adaptive multichannel LMS algorithm for blind SIMO identification
abstract
Adaptive algorithms for blindly identifying single-input multiple-output (SIMO) systems are appealing because of their computational efficiency and capability of continuously tracking a time-varying system. Adaptive multichannel least-mean-square (MCLMS) algorithms (with and without the unit-norm constraint) are analyzed, and the optimal step size is derived. A simple yet effective variable step-size MCLMS algorithm is proposed, and its performance is evaluated with simulations.
Yiteng Huang, Jacob Benesty, Jingdong Chen
IEEE Signal Process. Lett.1
2005 A Blind Channel Identification-Based Two-Stage Approach to Separation and Dereverberation of Speech Signals in a Reverberant Environment
abstract
Blind separation of independent speech sources from their convolutive mixtures in a reverberant acoustic environment is a difficult problem and the state-of-the-art blind source separation techniques are still unsatisfactory. The challenge lies in the coexistence of spatial interference from competing sources and temporal echoes due to room reverberation in the observed mixtures. Focusing only on optimizing the signal-to-interference ratio is inadequate for most if not all speech processing systems. In this paper, we deduce that spatial interference and temporal echoes can be separated and an M/spl times/N MIMO system will be converted into M SIMO systems that are free of spatial interference. Furthermore we show that the channel matrices of these SIMO systems are irreducible if the channels from the same source in the MIMO system do not share common zeros. Thereafter we can apply the Bezout theorem to remove reverberation in those SIMO systems. Such a two-stage procedure leads to a novel sequential source separation and speech dereverberation algorithm based on blind multichannel identification. Simulations with measurements obtained in the varechoic chamber at Bell Labs demonstrate the success and robustness of the proposed algorithm in highly reverberant acoustic environments.
Yiteng Huang, Jacob Benesty, Jingdong Chen
IEEE Trans. Speech Audio Process.1
2004 An exponentiated gradient adaptive algorithm for blind identification of sparse SIMO systems
abstract
Sparse impulse responses are encountered in many acoustic and wireless channels. Recently, a class of exponentiated gradient (EG) algorithms has been proposed. One of the algorithms belonging to this class, the so-called EG/spl plusmn/ algorithm, converges and tracks much better than the classical stochastic gradient, or LMS, algorithm for sparse impulse responses. We apply this technique to blind identification of a sparse SIMO system and develop the multichannel EG/spl plusmn/ algorithm. A simple experiment demonstrates its advantage in convergence compared to the MCLMS algorithm.
Jacob Benesty, Yiteng Huang, Jingdong Chen
ICASSP (2)2
2004 An adaptive blind SIMO identification approach to joint multichannel time delay estimation
abstract
Time delay estimation (TDE) is a difficult problem in a reverberant environment and the traditional generalized cross-correlation (GCC) methods perform poorly. The adaptive eigenvalue decomposition (AED) algorithm recently proposed by the authors exploits a blind channel identification (BCI) technique and deals with room reverberation more effectively. The AED algorithm was developed for a two-channel system. It requires that the two channels do not share any common zeros (a necessary condition of system identifiability). We generalize the AED algorithm to multichannel (more than 2) systems. Compared to the AED algorithm, the generalized method is more robust since it is less likely for all channels to share a common zero when more sensors are used.
Jingdong Chen, Yiteng Huang, Jacob Benesty
ICASSP (4)2
2004 Separating ISI and CCI in a two-step FIR Bezout equalizer for MIMO systems of frequency-selective channels
abstract
The use of multiple antennas at both the transmitter and the receiver in wireless communications implies a great channel capacity, but the signal detection is a crucial problem in achieving channel capacity, particularly in the general case of frequency selective channels, where both intersymbol interference (ISI) and cochannel interference (CCI) are significant. We show that ISI and CCI can be separated and then be cancelled in two different steps. We develop a two-step FIR Bezout equalizer and deduce the theoretically smallest length of its equalization filter.
Yiteng Huang, Jacob Benesty, Jingdong Chen
ICASSP (4)1
2004 Recognition of noisy speech using dynamic spectral subband centroids
abstract
Despite their widespread popularity as front-end parameters for speech recognition, the cepstral coefficients derived from either linear prediction analysis or a filter-bank are found to be sensitive to additive noise. In this letter, we discuss the use of spectral subband centroids for robust speech recognition. We show that centroids, if properly selected, can achieve recognition performance comparable to that of the mel-frequency cepstral coefficients (MFCCs) in clean speech, while delivering better performance than MFCC in noisy environments. A procedure is proposed to construct the dynamic centroid feature vector that essentially embodies the transitional spectral information. We discuss some properties of the proposed dynamic features.
Jingdong Chen, Yiteng Huang, Kuldip K. Paliwal
IEEE Signal Process. Lett.2
2004 Time-delay estimation via linear interpolation and cross correlation
abstract
Time-delay estimation (TDE), which aims at measuring the relative time difference of arrival (TDOA) between different channels is a fundamental approach for identifying, localizing, and tracking radiating sources. Recently, there has been a growing interest in the use of TDE based locator for applications such as automatic camera steering in a room conferencing environment where microphone sensors receive not only the direct-path signal, but also attenuated and delayed replicas of the source signal due to reflections from boundaries and objects in the room. This multipath propagation effect introduces echoes and spectral distortions into the observation signal, termed as reverberation, which severely deteriorates a TDE algorithm in its performance. This paper deals with the TDE problem with emphasis on combating reverberation using multiple microphone sensors. The multichannel cross correlation coefficient (MCCC) is rederived here, in a new way, to connect it to the well-known linear interpolation technique. Some interesting properties and bounds of the MCCC are discussed and a recursive algorithm is introduced so that the MCCC can be estimated and updated efficiently when new data snapshots are available. We then apply the MCCC to the TDE problem. The resulting new algorithm can be treated as a natural generalization of the generalized cross correlation (GCC) TDE method to the multichannel case. It is shown that this new algorithm can take advantage of the redundancy provided by multiple microphone sensors to improve TDE against both reverberation and noise. Experiments confirm that the relative time-delay estimation accuracy increases with the number of sensors.
Jacob Benesty, Jingdong Chen, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2003 A fast recursive algorithm for optimum sequential signal detection in a BLAST system
abstract
BLAST (Bell Laboratories layered Space-Time) wireless systems are multiple-antenna communication schemes which can achieve very high spectral efficiencies in scattering environments, with no increase in bandwidth or transmitted power. The most popular and, by far, the most practical architecture is the so-called vertical BLAST (V-BLAST). The signal detection algorithm of a V-BLAST system is computationally very intensive. If the number of transmitters is M and is equal to the number of receivers, this complexity is proportional to M/sup 4/ at each sample time. In this paper, we propose a simple and very efficient algorithm that reduces the complexity by a factor of M.
Jacob Benesty, Yiteng Huang, Jingdong Chen
ICASSP (5)2
2003 Robust time delay estimation exploiting spatial correlation
abstract
To find the position of an acoustic source in a room, a set of relative delays among different microphone pairs has to be determined. The generalized cross-correlation method is the most popular to do so and is well explained in a landmark paper (Knapp and Carter (1996)). In this paper, we show how we can take advantage of the redundancy when more than two microphones are available. It is believed that the redundancy will help to better cope with noise and reverberation. The idea of cross-correlation coefficient between two signals is generalized to the multichannel case by using the notion of spatial prediction. The multichannel spatial correlation matrix is then deduced and it is shown how it can be used for time delay estimation.
Jingdong Chen, Jacob Benesty, Yiteng Huang
ICASSP (5)3
2003 Adaptive blind identification of SIMO systems using channel cross-relation in the frequency domain
abstract
The implementation of existing methods for blind identification of single-input multiple-output (SIMO) systems is limited in practice since they are difficult to execute in an adaptive mode and are, in general, computationally intensive. We extend our previous study (Huang, Y. and Benesty, J., Sig. Processing, vol.82, no.8, p.99-110, 2002) into the frequency domain and propose an unconstrained normalized multi-channel frequency-domain LMS (UNMCFLMS) algorithm. Numerical simulations show that the UNMCFLMS algorithm performs as well as (for a SIMO system with relatively short channel impulse responses) or better than (for a SIMO system with long channel impulse responses) its time-domain counterpart and the cross-relation (CR) batch method in practical situations.
Yiteng Huang, Jacob Benesty, Jingdong Chen
ICASSP (6)1
2003 Robust time delay estimation exploiting redundancy among multiple microphones
abstract
To find the position of an acoustic source in a room, typically, a set of relative delays among different microphone pairs needs to be determined. The generalized cross-correlation (GCC) method is the most popular to do so and is well explained in a landmark paper by Knapp and Carter. In this paper, the idea of cross-correlation coefficient between two random signals is generalized to the multichannel case by using the notion of spatial prediction. The multichannel spatial correlation matrix is then deduced and its properties are discussed. We then propose a new method based on the multichannel spatial correlation matrix for time delay estimation. It is shown that this new approach can take advantage of the redundancy when more than two microphones are available and this redundancy can help the estimator to better cope with noise and reverberation.
Jingdong Chen, Jacob Benesty, Yiteng Huang
IEEE Trans. Speech Audio Process.3
2002 Adaptive blind channel identification: Multi-channel least mean square and Newton algorithms
abstract
The problem of identifying a single-input multiple-output FIR system without a training signal, the so-called blind system identification, is addressed and two adaptive multi-channel approaches, least mean square (LMS) and Newton algorithms, are proposed. In contrast to the existing batch blind channel identification schemes, the proposed algorithms construct an error signal based on the cross relations between different channels in a novel, systematic way. The corresponding cost (error) function is easy to manipulate and facilitates the use of adaptive filtering methods for an efficient blind channel identification scheme. It is theoretically shown and practically demonstrated by numerical studies that the proposed algorithms converge in the mean to the desired channel impulse responses for an identifiable system.
Yiteng Huang, Jacob Benesty
ICASSP1
2002 Recognition of noisy speech using normalized moments
Jingdong Chen, Yiteng Huang, Frank K. Soong
INTERSPEECH2
2002 Adaptive multi-channel least mean square and Newton algorithms for blind channel identification
Yiteng Huang, Jacob Benesty
Signal Process.1
2001 Real-time passive source localization: a practical linear-correction least-squares approach
abstract
A linear-correction least-squares estimation procedure is proposed for the source localization problem under an additive measurement error model. The method, which can be easily implemented in a real-time system with moderate computational complexity, yields an efficient source location estimator without assuming a priori knowledge of noise distribution. Alternative existing estimators, including likelihood-based, spherical intersection, spherical interpolation, and quadratic-correction least-squares estimators, are reviewed and comparisons of their complexity, estimation consistency and efficiency against the Cramer-Rao lower bound are made. Numerical studies demonstrate that the proposed estimator performs better under many practical situations.
Yiteng Huang, Jacob Benesty, Gary W. Elko, Russell M. Mersereati
IEEE Trans. Speech Audio Process.1
2000 Passive acoustic source localization for video camera steering
abstract
A multi-input one-step least-squares (OSLS) algorithm for passive source localization is proposed. It is shown that the OSLS algorithm is mathematically equivalent to the so-called spherical interpolation (SI) method but with less computational complexity. The OSLS/SI method uses spherical equations (instead of hyperbolic equations) and solves them in a least-squares sense. Based on the adaptive eigenvalue decomposition time delay estimation method previously proposed by the same authors and the OSLS source localization algorithm, a real-time passive source localization system for video camera steering is presented. The system demonstrates many desirable features such as accuracy, portability, and robustness.
Yiteng Huang, Jacob Benesty, Gary W. Elko
ICASSP1
1999 Adaptive eigenvalue decomposition algorithm for real time acoustic source localization system
abstract
To locate an acoustic source in a room, the relative delay between microphone pairs must be determined efficiently and accurately. However, most traditional time delay estimation (TDE) algorithms fail in reverberant environments. A new approach is proposed that takes into account the reverberation of the room. A real time PC-based TDE system running under Microsoft/sup TM/ Windows system was developed with three TDE techniques: classical cross-correlation, phase transform, and a new algorithm that is proposed in this paper. The system provides an interactive platform that allows users to compare performance of these algorithms.
Yiteng Huang, Jacob Benesty, Gary W. Elko
ICASSP1