EDBT 2026 Demo / reviewers in the wild / expert
Ke Tan 0001
dblp:56/5686-1
· DBLP profile ↗
30ranked-venue papers
14as first author
18since 2021 · last 2025
0000-0001-5073-8060ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 14 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder RefinementabstractDeploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their computational cost limits their feasibility on embedded platforms. This work presents an efficient end-to-end SE framework that leverages a Differentiable Digital Signal Processing (DDSP) vocoder for high-quality speech synthesis. First, a compact neural network predicts enhanced acoustic features from noisy speech: spectral envelope, fundamental frequency (F0), and periodicity. These features are fed into the DDSP vocoder to synthesize the enhanced waveform. The system is trained end-to-end with STFT and adversarial losses, enabling direct optimization at the feature and waveform levels. Experimental results show that our method improves intelligibility and quality by 4% (STOI) and 19% (DNSMOS) over strong baselines without significantly increasing computation, making it well-suited for real-time applications. Heitor R. Guimarães, Ke Tan 0001, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey 0004, Buye Xu |
ASRU | 2 |
| 2025 | Reexamining the Efficacy of MetricGAN for Speech EnhancementabstractMetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluation exclusively at high SNR. Firstly, we comprehensively assess MetricGAN models using mainstream metrics, surprisingly revealing MetricGAN models produce worse SISDR and STOI than unprocessed noisy speech. Secondly, we demonstrate that training MetricGAN models at low SNR often results in convergence to biased local minima, where PESQ scores are inflated while their SISDR and STOI values deteriorate significantly. In addition, we propose and validate two training tricks to address these issues: SISDR regularization and mixture-of-actor training. We find that these tricks effectively guide MetricGAN models to avoid local minima, thus improving speech quality. Ali Aroudi, Buye Xu, Ashutosh Pandey 0004, Francesco Nesta, Anurag Kumar 0003, Alexander Reich, Ke Tan 0001 |
ICASSP | 8 |
| 2024 | Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency DomainabstractWe introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual fusion, effectively merging audio and visual features and enabling seamless information interchange between auditory and visual modalities. These fused features drive a separator module, which separates the acoustic features of individual speakers. The separator module is based on the recently proposed TF-Gridnet, which comprises an intra-frame full-band component, a sub-band temporal module that captures frequency-specific temporal dependencies, and a cross-attention module dedicated to extracting long-term fused audiovisual features. To encourage the utilization of visual streams during training, we employ a Signal-to-Noise Ratio (SNR) scheduler. Experimental results demonstrate that the proposed model advances the state-of- the-art speaker separation performance in several audiovisual benchmark datasets. Vahid Ahmadi Kalkhorani, Anurag Kumar 0003, Ke Tan 0001, Buye Xu, DeLiang Wang |
ICASSP | 3 |
| 2024 | A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech EnhancementabstractSelf-supervised learned models have been found to be very effective for tasks such as automatic speech recognition, speaker identification, and others. However, their utility in speech enhancement systems is yet to be firmly established, and perhaps slightly misunderstood. In this paper, we investigate the uses of SSL representations for single-channel speech enhancement in challenging conditions and establish the impact they can have on the enhancement task. Our constraints are designed around on-device real-time speech enhancement – model being causal, and the compute footprint being small. Additionally, we focus on low SNR conditions where such models struggle to provide good performance. Ke Tan 0001, Buye Xu, Anurag Kumar 0003 |
ICASSP | 2 |
| 2024 | FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, Ke Tan 0001, Ashutosh Pandey 0004, Jung-Suk Lee, Buye Xu, Francesco Nesta |
INTERSPEECH | 3 |
| 2023 | Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech EnhancementabstractMost speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivariate Gaussian negative log-likelihood (NLL) improves SE performance at no extra cost. During training, our approach augments a model learning complex spectral mapping with a temporary submodel to predict the covariance of the enhancement error at each time-frequency bin. Due to unrestricted heteroscedas-tic uncertainty, the covariance introduces an undersampling effect, detrimental to SE performance. To mitigate undersampling, our approach inflates the uncertainty lower bound and weights each loss component with their uncertainty, effectively compensating severely undersampled components with more penalties. Our multivariate setting reveals common covariance assumptions such as scalar and diagonal matrices. By weakening these assumptions, we show that the NLL achieves superior performance compared to popular loss functions including the mean squared error (MSE), mean absolute error (MAE), and scale-invariant signal-to-distortion ratio (SI-SDR). Kuan-Lin Chen 0002, Daniel D. E. Wong, Ke Tan 0001, Buye Xu, Anurag Kumar 0003, Vamsi K. Ithapu |
ICASSP | 3 |
| 2023 | Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in TorchaudioabstractMeasuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a set of models to estimate such known metrics using deep neural networks. These models are made available in the well-established TorchAudio library, the core audio and speech processing library within the PyTorch deep learning framework. We refer to it as TorchAudio-Squim, TorchAudio-Speech QUality and Intelligibility Measures. More specifically, in the current version of TorchAudio-squim, we establish and release models for estimating PESQ, STOI and SI-SDR among objective metrics and MOS among subjective metrics. We develop a novel approach for objective metric estimation and use a recently developed approach for subjective metric estimation. These models operate in a "referenceless" manner, that is they do not require the corresponding clean speech as reference for speech assessment. Given the unavailability of clean speech and the effortful process of subjective evaluation in real-world situations, such easy-to-use tools would greatly benefit speech processing research and development. Anurag Kumar 0003, Ke Tan 0001, Zhaoheng Ni, Pranay Manocha, Xiaohui Zhang 0007, Ethan Henderson, Buye Xu |
ICASSP | 2 |
| 2023 | A Simple RNN Model for Lightweight, Low-compute and Low-latency Multichannel Speech Enhancement in the Time Domain
Ashutosh Pandey 0004, Ke Tan 0001, Buye Xu |
INTERSPEECH | 2 |
| 2023 | Time-domain Transformer-based Audiovisual Speaker Separation
Vahid Ahmadi Kalkhorani, Anurag Kumar 0003, Ke Tan 0001, Buye Xu, DeLiang Wang |
INTERSPEECH | 3 |
| 2023 | Rethinking Complex-Valued Deep Neural Networks for Monaural Speech Enhancement
Ke Tan 0001, Buye Xu, Anurag Kumar 0003 |
INTERSPEECH | 2 |
| 2022 | Location-Based Training for Multi-Channel Talker-Independent Speaker SeparationabstractPermutation-invariant training (PIT) is a dominant approach for addressing the permutation ambiguity problem in talker-independent speaker separation. Leveraging spatial information afforded by microphone arrays, we propose a new training approach to resolving permutation ambiguities for multi-channel speaker separation. The proposed approach, named location-based training (LBT), assigns speakers on the basis of their spatial locations. This training strategy is easy to apply, and organizes speakers according to their positions in physical space. Specifically, this study investigates azimuth angles and source distances for location-based training. Evaluation results on separating two- and three-speaker mixtures show that azimuth-based training consistently outperforms PIT, and distance-based training further improves the separation performance when speaker azimuths are close. Furthermore, we dynamically select azimuth-based or distance-based training by estimating the azimuths of separated speakers, which further improves separation performance. LBT has a linear training complexity with respect to the number of speakers, as opposed to the factorial complexity of PIT. We further demonstrate the effectiveness of LBT for the separation of four and five concurrent speakers. Hassan Taherian, Ke Tan 0001, DeLiang Wang |
ICASSP | 2 |
| 2022 | Multi-Channel Talker-Independent Speaker Separation Through Location-Based TrainingabstractPermutation ambiguity is a crucial issue for deep learning based talker-independent speaker separation. Deep clustering and permutation invariant training (PIT) have been widely used to address the permutation ambiguity problem in monaural scenarios. Although both approaches have been extended to multi-microphone scenarios, we believe that the permutation ambiguity problem can be naturally avoided by leveraging the spatial relations of multiple speakers. In this study, we present location-based training (LBT), a new approach to achieve talker independency in multi-channel speaker separation. Unlike PIT that examines all possible permutations, LBT assigns speakers according to their positions in physical space. With a linear training complexity to the number of concurrent speakers, LBT is computationally much more efficient than PIT with a factorial complexity, particularly when a large number of overlapping speakers needs to be separated. Specifically, we propose two training criteria: azimuth-based and distance-based training, using speaker azimuths and distances relative to a microphone array, respectively. Evaluation results show that LBT significantly outperforms PIT on two-speaker and three-speaker mixtures with different array geometries and in various acoustic conditions. In addition, we propose a joint training strategy to integrate azimuth-based and distance-based training, which further improves separation performance. Hassan Taherian, Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Neural Spectrospatial FilteringabstractAs the most widely-used spatial filtering approach for multi-channel speech separation, beamforming extracts the target speech signal arriving from a specific direction. An emerging alternative approach is multi-channel complex spectral mapping, which trains a deep neural network (DNN) to directly estimate the real and imaginary spectrograms of the target speech signal from those of the multi-channel noisy mixture. In this all-neural approach, the trained DNN itself becomes a nonlinear, time-varying spectrospatial filter. However, it remains unclear how this approach performs relative to commonly-used beamforming techniques on different array configurations and acoustic environments. This paper is devoted to examining this issue in a systematic way. Comprehensive evaluations show that multi-channel complex spectral mapping achieves separation performance comparable to or better than beamforming for different array geometries and speech separation tasks and reduces to monaural complex spectral mapping in single-channel conditions, demonstrating the general utility of this approach on multi-channel and single-channel speech separation. In addition, such an approach is computationally more efficient than widely-used mask-based beamforming. We conclude that this neural spectrospatial filter provides a strong alternative to traditional and mask-based beamforming. Ke Tan 0001, Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Compressing Deep Neural Networks for Efficient Speech EnhancementabstractThe use of deep neural networks (DNNs) has dramatically improved the performance of speech enhancement in the past decade. However, a large DNN is typically required to achieve strong enhancement performance, and this kind of model is both computationally intensive and memory consuming. Hence it is difficult to deploy such DNNs on devices with limited hardware resources or in applications with strict latency requirements. In order to address this problem, we propose a model compression pipeline to reduce DNN size for speech enhancement, which is based on three kinds of techniques: sparse regularization, iterative pruning and clustering-based quantization. Evaluation results show that our approach substantially reduces the sizes of different DNNs without significantly affecting their enhancement performance. Moreover, we find that training and compressing a large DNN yields higher STOI and PESQ than directly training a small DNN that has a comparable size to the compressed DNN. This further suggests the benefits of using the proposed model compression approach. Ke Tan 0001, DeLiang Wang |
ICASSP | 1 |
| 2021 | Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral MappingabstractSpeech quality and intelligibility can be severely degraded by back-ground noise in mobile communication. In order to attenuate back-ground noise, speech enhancement systems have been integrated into mobile phones, and a microphone array is typically deployed to improve the enhancement performance. This paper proposes a novel approach to real-time speech enhancement for dual-microphone mobile phones. Our approach employs a causal densely-connected convolutional recurrent network to perform dual-channel complex spectral mapping. We apply a structured pruning technique for compressing the model without significantly affecting the enhancement performance. This leads to a real-time enhancement system for on-device processing. Evaluation results show that the pro-posed approach substantially advances the performance of an earlier approach to dual-channel speech enhancement for mobile communication. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 1 |
| 2021 | SAGRNN: Self-Attentive Gated RNN For Binaural Speaker Separation With Interaural Cue PreservationabstractMost existing deep learning based binaural speaker separation systems focus on producing a monaural estimate for each of the target speakers, and thus do not preserve the interaural cues, which are crucial for human listeners to perform sound localization and lateralization. In this study, we address talker-independent binaural speaker separation with interaural cues preserved in the estimated binaural signals. Specifically, we extend a newly-developed gated recurrent neural network for monaural separation by additionally incorporating self-attention mechanisms and dense connectivity. We develop an end-to-end multiple-input multiple-output system, which directly maps from the binaural waveform of the mixture to those of the speech signals. The experimental results show that our proposed approach achieves significantly better separation performance than a recent binaural separation approach. In addition, our approach effectively preserves the interaural cues, which improves the accuracy of sound localization. Ke Tan 0001, Buye Xu, Anurag Kumar 0003, Eliya Nachmani, Yossi Adi |
IEEE Signal Process. Lett. | 1 |
| 2021 | Towards Model Compression for Deep Learning Based Speech EnhancementabstractThe use of deep neural networks (DNNs) has dramatically elevated the performance of speech enhancement over the last decade. However, to achieve strong enhancement performance typically requires a large DNN, which is both memory and computation consuming, making it difficult to deploy such speech enhancement systems on devices with limited hardware resources or in applications with strict latency requirements. In this study, we propose two compression pipelines to reduce the model size for DNN-based speech enhancement, which incorporates three different techniques: sparse regularization, iterative pruning and clustering-based quantization. We systematically investigate these techniques and evaluate the proposed compression pipelines. Experimental results demonstrate that our approach reduces the sizes of four different models by large margins without significantly sacrificing their enhancement performance. In addition, we find that the proposed approach performs well on speaker separation, which further demonstrates the effectiveness of the approach for compressing speech separation models. Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Deep Learning Based Real-Time Speech Enhancement for Dual-Microphone Mobile PhonesabstractIn mobile speech communication, speech signals can be severely corrupted by background noise when the far-end talker is in a noisy acoustic environment. To suppress background noise, speech enhancement systems are typically integrated into mobile phones, in which one or more microphones are deployed. In this study, we propose a novel deep learning based approach to real-time speech enhancement for dual-microphone mobile phones. The proposed approach employs a new densely-connected convolutional recurrent network to perform dual-channel complex spectral mapping. We utilize a structured pruning technique to compress the model without significantly degrading the enhancement performance, which yields a low-latency and memory-efficient enhancement system for real-time processing. Experimental results suggest that the proposed approach consistently outperforms an earlier approach to dual-channel speech enhancement for mobile phone communication, as well as a deep learning based beamformer. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Improving Robustness of Deep Learning Based Monaural Speech Enhancement Against Processing ArtifactsabstractIn voice telecommunication, the intelligibility and quality of speech signals can be severely degraded by background noise if the speaker at the transmitting end talks in a noisy environment. Therefore, a speech enhancement system is typically integrated into the transmitter device or the receiver device. Without the knowledge of whether the other end is equipped with a speech enhancer, the transmitter and receiver devices can both process a speech signal with their speech enhancers. In this study, we find that enhancing a speech signal twice can dramatically degrade the enhancement performance. This is because the downstream speech enhancer is sensitive to the processing artifacts introduced by the upstream enhancer. We analyze this problem and propose a new training scheme for the downstream deep learning based speech enhancement model. Our experimental results show that the proposed training strategy substantially elevate the robustness of speech enhancers against artifacts induced by another speech enhancer. Ke Tan 0001, DeLiang Wang |
ICASSP | 1 |
| 2020 | Learning Complex Spectral Mapping With Gated Convolutional Recurrent Networks for Monaural Speech EnhancementabstractPhase is important for perceptual quality of speech. However, it seems intractable to directly estimate phase spectra through supervised learning due to their lack of spectrotemporal structure in it. Complex spectral mapping aims to estimate the real and imaginary spectrograms of clean speech from those of noisy speech, which simultaneously enhances magnitude and phase responses of speech. Inspired by multi-task learning, we propose a gated convolutional recurrent network (GCRN) for complex spectral mapping, which amounts to a causal system for monaural speech enhancement. Our experimental results suggest that the proposed GCRN substantially outperforms an existing convolutional neural network (CNN) for complex spectral mapping in terms of both objective speech intelligibility and quality. Moreover, the proposed approach yields significantly higher STOI and PESQ than magnitude spectral mapping and complex ratio masking. We also find that complex spectral mapping with the proposed GCRN provides an effective phase estimate. Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Bridging the Gap Between Monaural Speech Enhancement and Recognition With Distortion-Independent Acoustic ModelingabstractMonaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago. Although enhanced speech has been demonstrated to have better intelligibility and quality for human listeners, feeding it directly to automatic speech recognition (ASR) systems trained with noisy speech has not produced expected improvements in ASR performance. The lack of an enhancement benefit on recognition, or the gap between monaural speech enhancement and recognition, is often attributed to speech distortions introduced in the enhancement process. In this article, we analyze the distortion problem, compare different acoustic models, and investigate a distortion-independent training scheme for monaural speech recognition. Experimental results suggest that distortion-independent acoustic modeling is able to overcome the distortion problem. Such an acoustic model can also work with speech enhancement models different from the one used during training. Moreover, the models investigated in this paper outperform the previous best system on the CHiME-2 corpus. Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Complex Spectral Mapping with a Convolutional Recurrent Network for Monaural Speech EnhancementabstractPhase is important for perceptual quality in speech enhancement. However, it seems intractable to directly estimate phase spectrogram through supervised learning due to lack of clear structure in phase spectrogram. Complex spectral mapping aims to estimate the real and imaginary spectrograms of clean speech from those of noisy speech, which simultaneously enhances magnitude and phase responses of noisy speech. In this paper, we propose a new convolutional recurrent network (CRN) for complex spectral mapping, which leads to a causal system for noise- and speaker-independent speech enhancement. In terms of objective intelligibility and perceptual quality, the proposed CRN significantly outperforms an existing convolutional neural network (CNN) for complex spectral mapping, as well as a strong CRN for magnitude spectral mapping. We additionally incorporate a newly-developed group strategy to substantially reduce the number of trainable parameters and the computational cost without sacrificing performance. Ke Tan 0001, DeLiang Wang |
ICASSP | 1 |
| 2019 | Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk ScenariosabstractIn mobile speech communication, the quality and intelligibility of the received speech can be severely degraded by background noise if the far-end talker is in an adverse acoustic environment. Therefore, speech enhancement algorithms are typically integrated into mobile phones to remove background noise. In this paper, we propose a novel deep learning based framework for real-time speech enhancement on dual-microphone mobile phones in a close-talk scenario. It incorporates a convolutional recurrent network (CRN) with high computational efficiency. In addition, the framework amounts to a causal system, which is necessary for real-time processing on mobile phones. We find that the proposed approach consistently outperforms a deep neural network (DNN) based method, as well as two traditional methods for speech enhancement. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 1 |
| 2019 | Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric PerspectiveabstractThis study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture oftwo sources, with their magnitudes accurately estimated and under a geometric constraint, the absolute phase difference between each source and the mixture can be uniquely determined; in addition, the source phases at each time-frequency T - F unit can be narrowed down to only two candidates. To pick the right candidate, we propose three algorithms based on iterative phase reconstruction, group delay estimation, and phase-difference sign prediction. State-of-the-art results are obtained on the publicly available wsj0-2mix and 3 mix corpus. Zhongqiu Wang 0001, Ke Tan 0001, DeLiang Wang |
ICASSP | 2 |
| 2019 | Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic ModelingabstractMonaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago.Although enhanced speech has been demonstrated to have better intelligibility and quality for human listeners, feeding it directly to automatic speech recognition (ASR) systems trained with noisy speech has not produced expected improvements in ASR performance.The lack of an enhancement benefit on recognition, or the gap between monaural speech enhancement and recognition, is often attributed to speech distortions introduced in the enhancement process.In this study, we analyze the distortion problem, compare different acoustic models, and investigate a distortionindependent training scheme for monaural speech recognition.Experimental results suggest that distortion-independent acoustic modeling is able to overcome the distortion problem.Such an acoustic model can also work with speech enhancement models different from the one used during training.Moreover, the models investigated in this paper outperform the previous best system on the CHiME-2 corpus. Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2019 | Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distortions
Hao Zhang 0112, Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2019 | Gated Residual Networks With Dilated Convolutions for Monaural Speech EnhancementabstractFor supervised speech enhancement, contextual information is important for accurate mask estimation or spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term contexts for tracking a target speaker, we treat speech enhancement as a sequence-to-sequence mapping, and present a novel convolutional neural network (CNN) architecture for monaural speech enhancement. The key idea is to systematically aggregate contexts through dilated convolutions, which significantly expand receptive fields. The CNN model additionally incorporates gating mechanisms and residual learning. Our experimental results suggest that the proposed model generalizes well to untrained noises and untrained speakers. It consistently outperforms a DNN, a unidirectional long short-term memory (LSTM) model and a bidirectional LSTM model in terms of objective speech intelligibility and quality metrics. Moreover, the proposed model has far fewer parameters than DNN and LSTM models. Ke Tan 0001, Jitong Chen, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Gated Residual Networks with Dilated Convolutions for Supervised Speech SeparationabstractIn supervised speech separation, deep neural networks (DNNs) are typically employed to predict an ideal time-frequency (T-F) mask in order to remove background interference. However, the performance of DNNs is frequently degraded for untrained noises and speakers. Inspired by recent research on dilated convolutions for context aggregation, we propose a novel convolutional neural network (CNN) to deal with noise- and speaker-independent speech separation. The proposed model incorporates dilated convolutions, gating mechanisms and residual learning. We find that the proposed model consistently outperforms a state-of-the-art long short-term memory (LSTM) based model in terms of objective speech intelligibility and quality. Additionally, the proposed CNN is more computationally efficient than the LSTM model. Ke Tan 0001, Jitong Chen, DeLiang Wang |
ICASSP | 1 |
| 2018 | A Convolutional Recurrent Neural Network for Real-Time Speech Enhancement
Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 1 |
| 2018 | A Two-Stage Approach to Noisy Cochannel Speech Separation with Gated Residual Networks
Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 1 |