VLDB 2026 Research / reviewers in the wild / expert
Buye Xu
dblp:120/6776
· DBLP profile ↗
37ranked-venue papers
0as first author
32since 2021 · last 2026
0000-0002-3027-7567ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 30 since 2021Artificial intelligence and machine learning · 17 · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audiovisual speech enhancement and voice activity detection using generative and regressive visual features
Vahid Ahmadi Kalkhorani, Buye Xu, DeLiang Wang |
Comput. Speech Lang. | 3 |
| 2025 | Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder RefinementabstractDeploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their computational cost limits their feasibility on embedded platforms. This work presents an efficient end-to-end SE framework that leverages a Differentiable Digital Signal Processing (DDSP) vocoder for high-quality speech synthesis. First, a compact neural network predicts enhanced acoustic features from noisy speech: spectral envelope, fundamental frequency (F0), and periodicity. These features are fed into the DDSP vocoder to synthesize the enhanced waveform. The system is trained end-to-end with STFT and adversarial losses, enabling direct optimization at the feature and waveform levels. Experimental results show that our method improves intelligibility and quality by 4% (STOI) and 19% (DNSMOS) over strong baselines without significantly increasing computation, making it well-suited for real-time applications. Heitor R. Guimarães, Ke Tan 0001, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey 0004, Buye Xu |
ASRU | 7 |
| 2025 | Robust Frame-level Speaker Localization in Reverberant and Noisy Environments by Exploiting Phase Difference LossesabstractThis paper investigates robust speaker localization at the frame level on the basis of complex spectral mapping, which is capable of learning both the magnitude and phase of the target signal. Unlike prevailing deep learning methods for speaker localization, we perform MIMO (multi-input multi-output) based multi-channel speech enhancement first and then localize the enhanced speaker using weighted generalized cross correlation. In addition, we propose new multi-channel loss functions that incorporate phase differences in order to preserve inter-channel phase relations, which is key to accurate sound localization. Systematic evaluations using simulated and recorded room impulse responses demonstrate that the proposed model yields excellent frame-level speaker localization results in reverberant and noisy environments and outperforms related methods by a large margin, even surpassing their utterance-level results. Shanmukha Srinivas Battula, Hassan Taherian, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
ICASSP | 5 |
| 2025 | Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech EnhancementabstractDeep learning-based speech enhancement (SE) methods often face significant computational challenges when needing to meet low-latency requirements because of the increased number of frames to be processed. This paper introduces the SlowFast framework which aims to reduce computation costs specifically when low-latency enhancement is needed. The framework consists of a slow branch that analyzes the acoustic environment at a low frame rate, and a fast branch that performs SE in the time domain at the needed higher frame rate to match the required latency. Specifically, the fast branch employs a state space model where its state transition process is dynamically modulated by the slow branch. Experiments on a SE task with a 2 ms algorithmic latency requirement using the Voice Bank + Demand dataset show that our approach reduces computation cost by 70% compared to a baseline single-branch network with equivalent parameters, without compromising enhancement performance. Furthermore, by leveraging the SlowFast framework, we implemented a network that achieves an algorithmic latency of just 62.5 μs (one sample point at 16 kHz sample rate) with a computation cost of 100 M MACs/s, while scoring a PESQ-NB of 3.12 and SISNR of 16.62. Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Vamsi K. Ithapu, Shih-Chii Liu |
ICASSP | 3 |
| 2025 | Advancing Active Speaker Detection for Egocentric VideosabstractThis paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing the lip region in the visual input, and (ii) applying motion blur augmentation. These methods significantly enhance the model’s performance in handling the challenges typical of egocentric videos. We showcase the effectiveness of these techniques on a simple but efficient causal audio-visual model. The proposed model, named EgoASD, demonstrates state-of-the-art performance on the EasyCom dataset, beating the previous SOTA by 1.7% mean Average Precision (mAP) with a model 2.5 times smaller. Our ablations highlight the importance of visual input, motion blur augmentation, the pretraining method and the importance of temporal context. To demonstrate its applicability in the real world, we apply our model to audio-visual speaker diarization, outperforming other baselines on EasyCom. Jaesung Huh, Juan Azcarreta, Anurag Kumar 0003, Ashutosh Pandey 0004, Ali Aroudi, Daniel D. E. Wong, Francesco Nesta, Buye Xu, Jacob Donley |
ICASSP | 8 |
| 2025 | Reexamining the Efficacy of MetricGAN for Speech EnhancementabstractMetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluation exclusively at high SNR. Firstly, we comprehensively assess MetricGAN models using mainstream metrics, surprisingly revealing MetricGAN models produce worse SISDR and STOI than unprocessed noisy speech. Secondly, we demonstrate that training MetricGAN models at low SNR often results in convergence to biased local minima, where PESQ scores are inflated while their SISDR and STOI values deteriorate significantly. In addition, we propose and validate two training tricks to address these issues: SISDR regularization and mixture-of-actor training. We find that these tricks effectively guide MetricGAN models to avoid local minima, thus improving speech quality. Ali Aroudi, Buye Xu, Ashutosh Pandey 0004, Francesco Nesta, Anurag Kumar 0003, Alexander Reich, Ke Tan 0001 |
ICASSP | 3 |
| 2025 | A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
Yuan-Kuei Wu, Juan Azcarreta, Kashyap Patel, Buye Xu, Jung-Suk Lee, Sanha Lee, Ashutosh Pandey 0004 |
INTERSPEECH | 4 |
| 2025 | Online AV-CrossNet: a Causal and Efficient Audiovisual System for Speech Enhancement and Target Speaker Extraction
Vahid Ahmadi Kalkhorani, Buye Xu, DeLiang Wang |
INTERSPEECH | 3 |
| 2024 | Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech EnhancementabstractWe present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and temporal processing within deep neural network (DNN) layers. Inspired by frequency-dependent multichannel filtering, our spatial filtering process applies multiple trainable filters to each hidden unit across the spatial dimension, resulting in a multichannel output. The temporal processing is applied over a single-channel output stream from the spatial processing using a Long Short-Term Memory (LSTM) network. The output from the temporal processing stage is then further integrated into the spatial dimension through elementwise multiplication. This explicit separation of spatial and temporal processing results in a resource-efficient network design. Empirical findings from our experiments show that our proposed model significantly outperforms robust baseline models while demanding far fewer parameters and computations, while achieving an ultra-low algorithmic latency of just 2 milliseconds. Ashutosh Pandey 0004, Buye Xu |
ICASSP | 2 |
| 2024 | On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech EnhancementabstractWe introduce a time-domain framework for efficient multichannel speech enhancement, emphasizing low latency and computational efficiency. This framework incorporates two compact deep neural networks (DNNs) surrounding a multichannel neural Wiener filter (NWF). The first DNN enhances the speech signal to estimate NWF coefficients, while the second DNN refines the output from the NWF. The NWF, while conceptually similar to the traditional frequency-domain Wiener filter, undergoes a training process optimized for low-latency speech enhancement, involving fine-tuning of both analysis and synthesis transforms. Our research results illustrate that the NWF output, having minimal nonlinear distortions, attains performance levels akin to those of the first DNN, deviating from conventional Wiener filter paradigms. Training all components jointly outperforms sequential training, despite its simplicity. Consequently, this framework achieves superior performance with fewer parameters and reduced computational demands, making it a compelling solution for resource-efficient multichannel speech enhancement. Tsun-An Hsieh, Jacob Donley, Buye Xu, Ashutosh Pandey 0004 |
ICASSP | 4 |
| 2024 | Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency DomainabstractWe introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual fusion, effectively merging audio and visual features and enabling seamless information interchange between auditory and visual modalities. These fused features drive a separator module, which separates the acoustic features of individual speakers. The separator module is based on the recently proposed TF-Gridnet, which comprises an intra-frame full-band component, a sub-band temporal module that captures frequency-specific temporal dependencies, and a cross-attention module dedicated to extracting long-term fused audiovisual features. To encourage the utilization of visual streams during training, we employ a Signal-to-Noise Ratio (SNR) scheduler. Experimental results demonstrate that the proposed model advances the state-of- the-art speaker separation performance in several audiovisual benchmark datasets. Vahid Ahmadi Kalkhorani, Anurag Kumar 0003, Ke Tan 0001, Buye Xu, DeLiang Wang |
ICASSP | 4 |
| 2024 | A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech EnhancementabstractSelf-supervised learned models have been found to be very effective for tasks such as automatic speech recognition, speaker identification, and others. However, their utility in speech enhancement systems is yet to be firmly established, and perhaps slightly misunderstood. In this paper, we investigate the uses of SSL representations for single-channel speech enhancement in challenging conditions and establish the impact they can have on the enhancement task. Our constraints are designed around on-device real-time speech enhancement – model being causal, and the compute footprint being small. Additionally, we focus on low SNR conditions where such models struggle to provide good performance. Ke Tan 0001, Buye Xu, Anurag Kumar 0003 |
ICASSP | 3 |
| 2024 | Leveraging Sound Localization to Improve Continuous Speaker SeparationabstractContinuous speaker separation aims to separate overlapping speakers in real-world environments like meetings, but it often falls short in isolating speech segments of a single speaker. This leads to split signals that adversely affect downstream applications such as automatic speech recognition and speaker diarization. Existing solutions like speaker counting have limitations. This paper presents a novel multi-channel approach for continuous speaker separation based on multi-input multi-output (MIMO) complex spectral mapping. This MIMO approach enables robust speaker localization by preserving inter-channel phase relations. Speaker localization as a byproduct of the MIMO separation model is then used to identify single-talker frames and reduce speaker splitting. We demonstrate that this approach achieves superior frame-level sound localization. Systematic experiments on the LibriCSS dataset further show that the proposed approach outperforms other methods, advancing state-of-the-art speaker separation performance. Hassan Taherian, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
ICASSP | 4 |
| 2024 | Dynamic Gated Recurrent Neural Network for Compute-efficient Speech EnhancementabstractThis paper introduces a new Dynamic Gated Recurrent Neural Network (DG-RNN) for compute-efficient speech enhancement models running on resource-constrained hardware platforms.It leverages the slow evolution characteristic of RNN hidden states over steps, and updates only a selected set of neurons at each step by adding a newly proposed select gate to the RNN model.This select gate allows the computation cost of the conventional RNN to be reduced during network inference.As a realization of the DG-RNN, we further propose the Dynamic Gated Recurrent Unit (D-GRU) which does not require additional parameters.Test results obtained from several state-ofthe-art compute-efficient RNN-based speech enhancement architectures using the DNS challenge dataset, show that the D-GRU based model variants maintain similar speech intelligibility and quality metrics comparable to the baseline GRU based models even with an average 50% reduction in GRU computes. Longbiao Cheng, Ashutosh Pandey 0004, Buye Xu, Tobi Delbruck, Shih-Chii Liu |
INTERSPEECH | 3 |
| 2024 | All Neural Low-latency Directional Speech Extraction
Ashutosh Pandey 0004, Sanha Lee, Juan Azcarreta, Buye Xu |
INTERSPEECH | 5 |
| 2024 | Towards Explainable Monaural Speaker Separation with Auditory-based Training
Hassan Taherian, Vahid Ahmadi Kalkhorani, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
INTERSPEECH | 5 |
| 2024 | FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, Ke Tan 0001, Ashutosh Pandey 0004, Jung-Suk Lee, Buye Xu, Francesco Nesta |
INTERSPEECH | 6 |
| 2023 | Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech EnhancementabstractMost speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivariate Gaussian negative log-likelihood (NLL) improves SE performance at no extra cost. During training, our approach augments a model learning complex spectral mapping with a temporary submodel to predict the covariance of the enhancement error at each time-frequency bin. Due to unrestricted heteroscedas-tic uncertainty, the covariance introduces an undersampling effect, detrimental to SE performance. To mitigate undersampling, our approach inflates the uncertainty lower bound and weights each loss component with their uncertainty, effectively compensating severely undersampled components with more penalties. Our multivariate setting reveals common covariance assumptions such as scalar and diagonal matrices. By weakening these assumptions, we show that the NLL achieves superior performance compared to popular loss functions including the mean squared error (MSE), mean absolute error (MAE), and scale-invariant signal-to-distortion ratio (SI-SDR). Kuan-Lin Chen 0002, Daniel D. E. Wong, Ke Tan 0001, Buye Xu, Anurag Kumar 0003, Vamsi K. Ithapu |
ICASSP | 4 |
| 2023 | Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in TorchaudioabstractMeasuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a set of models to estimate such known metrics using deep neural networks. These models are made available in the well-established TorchAudio library, the core audio and speech processing library within the PyTorch deep learning framework. We refer to it as TorchAudio-Squim, TorchAudio-Speech QUality and Intelligibility Measures. More specifically, in the current version of TorchAudio-squim, we establish and release models for estimating PESQ, STOI and SI-SDR among objective metrics and MOS among subjective metrics. We develop a novel approach for objective metric estimation and use a recently developed approach for subjective metric estimation. These models operate in a "referenceless" manner, that is they do not require the corresponding clean speech as reference for speech assessment. Given the unavailability of clean speech and the effortful process of subjective evaluation in real-world situations, such easy-to-use tools would greatly benefit speech processing research and development. Anurag Kumar 0003, Ke Tan 0001, Zhaoheng Ni, Pranay Manocha, Xiaohui Zhang 0007, Ethan Henderson, Buye Xu |
ICASSP | 7 |
| 2023 | LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural VocodersabstractAudio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker’s lip movements. This approach has been shown to yield improvements over audio-only speech enhancement, particularly for the removal of interfering speech. Despite recent advances in speech synthesis, most audio-visual approaches continue to use spectral mapping/masking to reproduce the clean audio, often resulting in visual backbones added to existing speech enhancement architectures. In this work, we propose LA-VocE, a new two-stage approach that predicts mel-spectrograms from noisy audio-visual speech via a transformer-based architecture, and then converts them into waveform audio using a neural vocoder (HiFi-GAN). We train and evaluate our framework on thousands of speakers and 11+ different languages, and study our model’s ability to adapt to different levels of background noise and speech interference. Our experiments show that LA-VocE outperforms existing methods according to multiple metrics, particularly under very noisy scenarios. Rodrigo Mira, Buye Xu, Jacob Donley, Anurag Kumar 0003, Stavros Petridis, Vamsi K. Ithapu, Maja Pantic |
ICASSP | 2 |
| 2023 | A Simple RNN Model for Lightweight, Low-compute and Low-latency Multichannel Speech Enhancement in the Time Domain
Ashutosh Pandey 0004, Ke Tan 0001, Buye Xu |
INTERSPEECH | 3 |
| 2023 | Time-domain Transformer-based Audiovisual Speaker Separation
Vahid Ahmadi Kalkhorani, Anurag Kumar 0003, Ke Tan 0001, Buye Xu, DeLiang Wang |
INTERSPEECH | 4 |
| 2023 | Multi-input Multi-output Complex Spectral Mapping for Speaker Separation
Hassan Taherian, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
INTERSPEECH | 4 |
| 2023 | Rethinking Complex-Valued Deep Neural Networks for Monaural Speech Enhancement
Ke Tan 0001, Buye Xu, Anurag Kumar 0003 |
INTERSPEECH | 3 |
| 2022 | TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech EnhancementabstractIn this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes speech signals from all channels independently using a dual-path attentive recurrent network (ARN), which is a recurrent neural network (RNN) augmented with self-attention. Next, an ARN is introduced along the spatial dimension for spatial context aggregation. TPARN is designed as a multiple-input and multiple-output architecture to enhance all input channels simultaneously. Experimental results demonstrate the superiority of TPARN over existing state-of-the-art approaches. Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang |
ICASSP | 2 |
| 2022 | Multichannel Speech Enhancement Without BeamformingabstractDeep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post-filter in a multistage processing provides additional improvements. In this work, we propose a two-stage strategy for multi-channel speech enhancement that does not require a traditional beamformer for additional performance. First, we propose a novel attentive dense convolutional network (ADCN) for estimating real and imaginary parts of complex spectrogram. ADCN obtains state-of-the-art results among single-stage models. Next, we use ADCN with a recently proposed triple-path attentive recurrent network (TPARN) for estimating waveform samples. The proposed strategy uses two insights; first, using different approaches in two stages; and second, using a stronger model in the first stage. We illustrate the efficacy of our strategy by evaluating multiple models in a two-stage approach with and without a traditional beamformer. Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang |
ICASSP | 2 |
| 2022 | Continual Self-Training With Bootstrapped Remixing For Speech EnhancementabstractWe propose RemixIT, a simple and novel self-supervised training method for speech enhancement. The proposed method is based on a continuously self-training scheme that overcomes limitations from previous studies including assumptions for the in-domain noise distribution and having access to clean target signals. Specifically, a separation teacher model is pre-trained on an out-of-domain dataset and is used to infer estimated target signals for a batch of in-domain mixtures. Next, we bootstrap the mixing process by generating artificial mixtures using permuted estimated clean and noise signals. Finally, the student model is trained using the permuted estimated sources as targets while we periodically update teacher’s weights using the latest student model. Our experiments show that RemixIT outperforms several previous state-of-the-art self-supervised methods under multiple speech enhancement tasks. Additionally, RemixIT provides a seamless alternative for semi-supervised and unsupervised domain adaptation for speech enhancement tasks, while being general enough to be applied to any separation task and paired with any separation model. Efthymios Tzinis, Yossi Adi, Vamsi K. Ithapu, Buye Xu, Anurag Kumar 0003 |
ICASSP | 4 |
| 2022 | Time-domain Ad-hoc Array Speech Enhancement Using a Triple-path NetworkabstractDeep neural networks (DNNs) are very effective for multichannel speech enhancement with fixed array geometries.However, it is not trivial to use DNNs for ad-hoc arrays with unknown order and placement of microphones.We propose a novel triplepath network for ad-hoc array processing in the time domain.The key idea in the network design is to divide the overall processing into spatial processing and temporal processing and use self-attention for spatial processing.Using self-attention for spatial processing makes the network invariant to the order and the number of microphones.The temporal processing is done independently for all channels using a recently proposed dual-path attentive recurrent network.The proposed network is a multiple-input multiple-output architecture that can simultaneously enhance signals at all microphones.Experimental results demonstrate the excellent performance of the proposed approach.Further, we present analysis to demonstrate the effectiveness of the proposed network in utilizing multichannel information even from microphones at far locations. Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang |
INTERSPEECH | 2 |
| 2022 | SAQAM: Spatial Audio Quality Assessment MetricabstractAudio quality assessment is critical for assessing the perceptual realism of sounds.However, the time and expense of obtaining "gold standard" human judgments limit the availability of such data.For AR&VR, good perceived sound quality and localizability of sources are among the key elements to ensure complete immersion of the user.Our work introduces SAQAM which uses a multi-task learning framework to assess listening quality (LQ) and spatialization quality (SQ) between any given pair of binaural signals without using any subjective data.We model LQ by training on a simulated dataset of triplet human judgments, and SQ by utilizing activation-level distances from networks trained for direction of arrival (DOA) estimation.We show that SAQAM correlates well with human responses across four diverse datasets.Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks.For example, simply replacing an existing loss with our metric yields improvement in a speech-enhancement network. Pranay Manocha, Anurag Kumar 0003, Buye Xu, Anjali Menon, Israel D. Gebru, Vamsi K. Ithapu, Paul Calamia |
INTERSPEECH | 3 |
| 2021 | Incorporating Real-World Noisy Speech in Neural-Network-Based Speech Enhancement SystemsabstractSupervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-world degraded speech data that may better represent the scenarios where such systems are used. In this paper, we explore methods that enable supervised speech enhancement systems to train on real-world degraded speech data. Specifically, we propose a semi-supervised approach for speech enhancement in which we first train a modified vector-quantized variational autoencoder that solves a source separation task. We then use this trained autoencoder to further train an enhancement network using real-world noisy speech data by computing a triplet-based unsupervised loss function. Experiments show promising results for incorporating real-world data in training speech enhancement systems. Yangyang Xia, Buye Xu, Anurag Kumar 0003 |
ASRU | 2 |
| 2021 | NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesabstractThe perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA. Clearly, these methods fail in real-world scenarios where the ground truth clean references are not available. In recent years, non-intrusive methods that train neural networks to predict ratings or scores have attracted much attention, but they suffer from several shortcomings such as lack of robustness, reliance on labeled data for training and so on. In this work, we propose a new direction for speech quality assessment. Inspired by human's innate ability to compare and assess the quality of speech signals even when they have non-matching contents, we propose a novel framework that predicts a subjective relative quality score for the given speech signal with respect to any provided reference without using any subjective data. We show that neural networks trained using our framework produce scores that correlate well with subjective mean opinion scores (MOS) and are also competitive to methods such as DNSMOS, which explicitly relies on MOS from humans for training networks. Moreover, our method also provides a natural way to embed quality-related information in neural networks, which we show is helpful for downstream tasks such as speech enhancement. Pranay Manocha, Buye Xu, Anurag Kumar 0003 |
NeurIPS | 2 |
| 2021 | SAGRNN: Self-Attentive Gated RNN For Binaural Speaker Separation With Interaural Cue PreservationabstractMost existing deep learning based binaural speaker separation systems focus on producing a monaural estimate for each of the target speakers, and thus do not preserve the interaural cues, which are crucial for human listeners to perform sound localization and lateralization. In this study, we address talker-independent binaural speaker separation with interaural cues preserved in the estimated binaural signals. Specifically, we extend a newly-developed gated recurrent neural network for monaural separation by additionally incorporating self-attention mechanisms and dense connectivity. We develop an end-to-end multiple-input multiple-output system, which directly maps from the binaural waveform of the mixture to those of the speech signals. The experimental results show that our proposed approach achieves significantly better separation performance than a recent binaural separation approach. In addition, our approach effectively preserves the interaural cues, which improves the accuracy of sound localization. Ke Tan 0001, Buye Xu, Anurag Kumar 0003, Eliya Nachmani, Yossi Adi |
IEEE Signal Process. Lett. | 2 |
| 2020 | Monaural Speech Dereverberation Using Temporal Convolutional Networks With Self AttentionabstractIn daily listening environments, human speech is often degraded by room reverberation, especially under highly reverberant conditions. Such degradation poses a challenge for many speech processing systems, where the performance becomes much worse than in anechoic environments. To combat the effect of reverberation, we propose a monaural (single-channel) speech dereverberation algorithm using temporal convolutional networks with self attention. Specifically, the proposed system includes a self-attention module to produce dynamic representations given input features, a temporal convolutional network to learn a nonlinear mapping from such representations to the magnitude spectrum of anechoic speech, and a one-dimensional (1-D) convolution module to smooth the enhanced magnitude among adjacent frames. Systematic evaluations demonstrate that the proposed algorithm improves objective metrics of speech quality in a wide range of reverberant conditions. In addition, it generalizes well to untrained reverberation times, room sizes, measured room impulse responses, real-world recorded noisy-reverberant speech, and different speakers. Yan Zhao 0010, DeLiang Wang, Buye Xu, Tao Zhang 0024 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term MemoryabstractHuman speech is usually distorted by room reverberation. These corruptions degrade speech quality and intelligibility, especially under a long reverberation time, and they also pose a serious problem for many speech-related applications such as automatic speech recognition. In this paper, we propose a supervised speech dereverberation algorithm that models late reverberation using a recurrent neural network (RNN) with long short-term memory (LSTM). By taking advantage of LSTM's ability to capture a long history, late reverberation can be effectively removed by the proposed approach. Systematic evaluations indicate that our approach improves the quality of reverberant speech in a wide range of reverberant conditions. Moreover, the proposed system is a causal system, which can be applied in real-time applications. Yan Zhao 0010, DeLiang Wang, Buye Xu, Tao Zhang 0024 |
ICASSP | 3 |
| 2018 | Perceptually Guided Speech Enhancement Using Deep Neural NetworksabstractHuman listeners often have difficulties understanding speech in the presence of background noise in the real world. Recently, supervised learning based speech enhancement approaches have achieved substantial success, and show significant improvements over the conventional approaches. However, existing supervised learning based approaches often try to minimize the mean squared error between the enhanced output and the pre-defined training target (e.g., the log power spectrum of clean speech), even though the purpose of such speech enhancement is to improve speech understanding in noise. In this paper, we propose a new deep neural networks based enhancement approach by incorporating a speech perception model into the loss function. Specifically, we use the short-time objective intelligibility metric in the loss in addition to the mean squared error. Optimizing the proposed perceptually guided loss is expected to improve speech intelligibility further. Systematic evaluations show that our proposed approach is able to improve speech intelligibility in a wide range of signal-to-noise ratios and noise types while maintaining speech quality. Yan Zhao 0010, Buye Xu, Ritwik Giri, Tao Zhang 0024 |
ICASSP | 2 |
| 2014 | Design of a high order binaural microphone array for hearing aids using a rigid spherical modelabstractWireless technology has allowed for a much wider variety in the design of microphone arrays for binaural hearing aids. To facilitate the design of these microphone arrays, this paper investigates the use of a spherical head model in the design of bilateral and binaural microphone arrays for hearing aids. The arrays have been designed using a free-field model, a spherical model, measurements on an artificial head, and measurements on an artificial head + torso. The results show that the free-field/spherical models over-estimate the speech-intelligibility weighted directivity index (SII-DI) of the bilateral and binaural arrays by respectively 0.9/0.4 and 0.8/0.5 dB. Furthermore the weights designed with the free-field/spherical model yield an SII-DI that is 0.7/0.6 dB lower for bilateral arrays and 0.9/0.9 dB lower for binaural arrays than the optimal SII-DI. Although the results show that the spherical model is better in predicting the DI than the free-field model, the spherical model does not design better weights. Ivo Merks, Buye Xu, Tao Zhang 0024 |
ICASSP | 2 |
| 2012 | Annoyance perception and modeling for hearing-impaired listenersabstractPerceptual annoyance of environmental sounds is measured for normal-hearing and hearing-impaired listeners under iso-level and iso-loudness conditions. Data from the hearing-impaired listeners shows similar trends to that from normal-hearing listeners, but with greater variability across individuals. A regression model based on the statistics of specific loudness and other perceptual features is fit to the data from the normal-hearing listeners, and is used to predict annoyance for the hearing-impaired listeners. Differences across the subject populations are discussed. Srikanth Vishnubhotla, Jinjun Xiao, Buye Xu, Martin F. McKinney, Tao Zhang 0024 |
ICASSP | 3 |