EDBT 2026 Demo / reviewers in the wild / expert
DeLiang Wang
dblp:31/6085 · also DeLiang L. Wang
· DBLP profile ↗
331ranked-venue papers
22as first author
67since 2021 · last 2026
0000-0001-8195-6319ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 200 · 18 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 176 · 2 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards decoupling frontend enhancement and backend recognition in monaural robust ASRabstractIt has been shown that the intelligibility of noisy speech can be improved by speech enhancement (SE) algorithms. However, monaural SE has not been established as an effective frontend for automatic speech recognition (ASR) in noisy conditions compared to an ASR model trained on noisy speech directly. The divide between SE and ASR impedes the progress of robust ASR systems, especially as SE has made major advances in recent years. This paper focuses on eliminating this divide with an ARN (attentive recurrent network) time-domain, a TF-CrossNet time-frequency domain, and an MP-SENet magnitude-phase based enhancement model. The proposed systems decouple frontend enhancement and backend ASR, with the latter trained only on clean speech. Results on the WSJ, CHiME-2, LibriSpeech, and CHiME-4 corpora demonstrate that ARN, TF-CrossNet, and MP-SENet enhanced speech all translate to improved ASR results in noisy and reverberant environments, and generalize well to real acoustic scenarios. The proposed system outperforms the baselines trained on corrupted speech directly. Furthermore, it cuts the previous best word error rate (WER) on CHiME-2 by 28.4% relatively with a 5.6% WER, and achieves 3.3/4.4% WER on single-channel CHiME-4 simulated/real test data without training on CHiME-4. We also observe consistent improvements using noise-robust Whisper as the backend ASR model. Ashutosh Pandey 0004, DeLiang Wang |
Comput. Speech Lang. | 3 |
| 2026 | Audiovisual speech enhancement and voice activity detection using generative and regressive visual features
Vahid Ahmadi Kalkhorani, Buye Xu, DeLiang Wang |
Comput. Speech Lang. | 4 |
| 2026 | Elevating robust multi-talker ASR by decoupling speaker separation and speech recognitionabstractDespite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing artifacts that degrade the ASR backend trained on clean speech. As a result, mainstream robust ASR systems train the backend on noisy speech to mitigate processing artifacts. In this work, we propose to decouple the training of the speaker separation frontend and the ASR backend, and evaluate the proposed system with the backend trained on clean speech only. Our decoupled system achieves 5.1% word error rates (WER) on the Libri2Mix dev/test sets, significantly outperforming other multi-talker ASR baselines. Its effectiveness is also demonstrated with the state-of-the-art 7.60%/5.74% WERs on 1-ch and 6-ch SMS-WSJ. Furthermore, on recorded LibriCSS, we achieve the speaker-attributed WER of 2.92%. These state-of-the-art results suggest that decoupling speaker separation and recognition is an effective approach to elevate robust multi-talker ASR performance. Finally, we provide insights into the acoustic conditions where the decoupled approach is expected to outperform the mainstream approach of training on noisy speech. Hassan Taherian, Vahid Ahmadi Kalkhorani, DeLiang Wang |
Speech Commun. | 4 |
| 2025 | Robust Frame-level Speaker Localization in Reverberant and Noisy Environments by Exploiting Phase Difference LossesabstractThis paper investigates robust speaker localization at the frame level on the basis of complex spectral mapping, which is capable of learning both the magnitude and phase of the target signal. Unlike prevailing deep learning methods for speaker localization, we perform MIMO (multi-input multi-output) based multi-channel speech enhancement first and then localize the enhanced speaker using weighted generalized cross correlation. In addition, we propose new multi-channel loss functions that incorporate phase differences in order to preserve inter-channel phase relations, which is key to accurate sound localization. Systematic evaluations using simulated and recorded room impulse responses demonstrate that the proposed model yields excellent frame-level speaker localization results in reverberant and noisy environments and outperforms related methods by a large margin, even surpassing their utterance-level results. Shanmukha Srinivas Battula, Hassan Taherian, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
ICASSP | 6 |
| 2025 | Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech RecognitionabstractDespite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing artifacts that degrade the ASR backend trained on clean speech. As a result, mainstream robust ASR systems train on noisy speech to avoid processing artifacts. In this work, we propose to decouple the training of the multi-channel speaker separation frontend and the ASR backend, with the latter trained only on clean speech. On SMS-WSJ, the proposed approach achieves a word error rate (WER) of 5.74%, outperforming the previous best by 14.3%. Furthermore, on recorded LibriCSS, we achieve the speaker-attributed WER of 3.86%, outperforming the previous best system trained on the same data by 24.8%. These state-of-the-art results suggest that decoupling speech separation and recognition is a potentially effective approach to robust ASR. Hassan Taherian, Vahid Ahmadi Kalkhorani, DeLiang Wang |
ICASSP | 4 |
| 2025 | Online AV-CrossNet: a Causal and Efficient Audiovisual System for Speech Enhancement and Target Speaker Extraction
Vahid Ahmadi Kalkhorani, Buye Xu, DeLiang Wang |
INTERSPEECH | 4 |
| 2025 | Combined generative and predictive modeling for speech super-resolution
Heming Wang, Eric W. Healy, DeLiang Wang |
Comput. Speech Lang. | 3 |
| 2025 | A systematic study of DNN based speech enhancement in reverberant and reverberant-noisy environments
Heming Wang, Ashutosh Pandey 0004, DeLiang Wang |
Comput. Speech Lang. | 3 |
| 2024 | Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency DomainabstractWe introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual fusion, effectively merging audio and visual features and enabling seamless information interchange between auditory and visual modalities. These fused features drive a separator module, which separates the acoustic features of individual speakers. The separator module is based on the recently proposed TF-Gridnet, which comprises an intra-frame full-band component, a sub-band temporal module that captures frequency-specific temporal dependencies, and a cross-attention module dedicated to extracting long-term fused audiovisual features. To encourage the utilization of visual streams during training, we employ a Signal-to-Noise Ratio (SNR) scheduler. Experimental results demonstrate that the proposed model advances the state-of- the-art speaker separation performance in several audiovisual benchmark datasets. Vahid Ahmadi Kalkhorani, Anurag Kumar 0003, Ke Tan 0001, Buye Xu, DeLiang Wang |
ICASSP | 5 |
| 2024 | Leveraging Sound Localization to Improve Continuous Speaker SeparationabstractContinuous speaker separation aims to separate overlapping speakers in real-world environments like meetings, but it often falls short in isolating speech segments of a single speaker. This leads to split signals that adversely affect downstream applications such as automatic speech recognition and speaker diarization. Existing solutions like speaker counting have limitations. This paper presents a novel multi-channel approach for continuous speaker separation based on multi-input multi-output (MIMO) complex spectral mapping. This MIMO approach enables robust speaker localization by preserving inter-channel phase relations. Speaker localization as a byproduct of the MIMO separation model is then used to identify single-talker frames and reduce speaker splitting. We demonstrate that this approach achieves superior frame-level sound localization. Systematic experiments on the LibriCSS dataset further show that the proposed approach outperforms other methods, advancing state-of-the-art speaker separation performance. Hassan Taherian, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
ICASSP | 5 |
| 2024 | Towards Explainable Monaural Speaker Separation with Auditory-based Training
Hassan Taherian, Vahid Ahmadi Kalkhorani, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
INTERSPEECH | 6 |
| 2024 | A surge of submissions
Taro Toyoizumi, DeLiang Wang |
Neural Networks | 2 |
| 2024 | Expansion of the editorial team
DeLiang Wang, Mauro Forti, Tongliang Liu, Taro Toyoizumi |
Neural Networks | 1 |
| 2024 | TF-CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker SeparationabstractWe introduce TF-CrossNet, a complex spectral mapping approach to speaker separation and enhancement in reverberant and noisy conditions. The proposed architecture comprises an encoder layer, a global multi-head self-attention module, a cross-band module, a narrow-band module, and an output layer. TF-CrossNet captures global, cross-band, and narrow-band correlations in the time-frequency domain. To address performance degradation in long utterances, we introduce a random chunk positional encoding. Experimental results on multiple datasets demonstrate the effectiveness and robustness of TF-CrossNet, achieving state-of-the-art performance in tasks including reverberant and noisy-reverberant speaker separation. Furthermore, TF-CrossNet exhibits faster and more stable training in comparison to recent baselines. Additionally, TF-CrossNet's high performance extends to multi-microphone conditions, demonstrating its versatility in various acoustic scenarios. Vahid Ahmadi Kalkhorani, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Multi-Channel Conversational Speaker Separation via Neural DiarizationabstractWhen dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or meeting environments, continuous speaker separation (CSS) is commonly employed. However, CSS requires a short separation window to avoid many speakers inside the window and sequential grouping of discontinuous speech segments. To address these limitations, we introduce a new multi-channel framework called “speaker separation via neural diarization” (SSND) for meeting environments. Our approach utilizes an end-to-end diarization system to identify the speech activity of each individual speaker. By leveraging estimated speaker boundaries, we generate a sequence of embeddings, which in turn facilitate the assignment of speakers to the outputs of a multi-talker separation model. SSND addresses the permutation ambiguity issue of talker-independent speaker separation during the diarization phase through location-based training, rather than during the separation process. This unique approach allows multiple non-overlapped speakers to be assigned to the same output stream, making it possible to efficiently process long segments-a task impossible with CSS. Additionally, SSND is naturally suitable for speaker-attributed ASR. We evaluate our proposed diarization and separation methods on the open LibriCSS dataset, advancing state-of-the-art diarization and ASR results by a large margin. Hassan Taherian, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Neuralkalman: A Learnable Kalman Filter for Acoustic Echo CancellationabstractThe robustness of the Kalman filter to double talk and its rapid convergence make it a popular approach for addressing acoustic echo cancellation (AEC) challenges. However, the inability to model nonlinearity and the need to tune control parameters cast limitations on such adaptive filtering algorithms. In this paper, we integrate the frequency domain Kalman filter (FDKF) and deep neural networks (DNNs) into a hybrid method, called NeuralKalman, to leverage the advantages of deep learning and adaptive filtering algorithms. Specifically, we employ a DNN to estimate nonlinearly distorted far-end signals, a transition factor, and the nonlinear transition function in the state equation of the FDKF algorithm. Experimental results show that the proposed NeuralKalman improves the performance of FDKF significantly and outperforms strong baseline methods. Yixuan Zhang 0005, Meng Yu 0003, Hao Zhang 0112, Dong Yu 0001, DeLiang Wang |
ASRU | 5 |
| 2023 | Multi-Resolution Location-Based Training for Multi-Channel Continuous Speech SeparationabstractThe performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently, location-based training (LBT) was proposed as a new training criterion for multi-channel talker-independent speaker separation. Assuming fixed array geometry, LBT outperforms widely-used permutation-invariant training in fully overlapped utterances and matched reverberant conditions. This paper extends LBT to conversational multi-channel speaker separation. We introduce multi-resolution LBT to estimate the complex spectrograms from low to high time and frequency resolutions. With multi-resolution LBT, convolutional kernels are assigned consistently based on speaker locations in physical space. Evaluation results show that multi-resolution LBT consistently outperforms other competitive methods on the recorded LibriCSS corpus. Hassan Taherian, DeLiang Wang |
ICASSP | 2 |
| 2023 | DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation TasksabstractSelf-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem, in this paper, we propose data2vec-SG (Speech Generation), which is a teacher-student learning framework that addresses speech generation tasks. Our data2vec-SG introduces a reconstruction module into data2vec [1] and enforces the representations to contain not only the semantic information but also the acoustic knowledge to generate clean speech waveforms. Experimental results demonstrate that the proposed framework boosts the performance of various speech generation tasks including speech enhancement, speech separation, and packet loss concealment. Meanwhile, the learned representation is also capable of helping other downstream tasks, which is demonstrated by the good performance in the speech recognition task in both clean and noisy conditions. Heming Wang, Yao Qian, Hemin Yang, Naoyuki Kanda, Takuya Yoshioka, Xiaofei Wang 0009, Shujie Liu 0001, Zhuo Chen 0006, DeLiang Wang, Michael Zeng 0001 |
ICASSP | 11 |
| 2023 | Cross-Domain Diffusion Based Speech Enhancement for Very Noisy SpeechabstractDeep learning based speech enhancement has achieved remarkable success, but challenges remain in low signal-to-noise ratio (SNR) nonstationary noise scenarios. In this study, we propose to incorporate diffusion-based learning into an enhancement model and improve robustness in extremely noisy conditions. Specifically, a frequency-domain diffusion-based generative module is employed, and it accepts the enhanced signal obtained from a time-domain supervised enhancement module as an auxiliary input to learn to recover clean speech spectrograms. Experimental results on the TIMIT dataset demonstrate the advantage of this approach and show better enhancement performance over other strong baselines in both -5 and -10 dB SNR noisy conditions. Heming Wang, DeLiang Wang |
ICASSP | 2 |
| 2023 | Time-domain Transformer-based Audiovisual Speaker Separation
Vahid Ahmadi Kalkhorani, Anurag Kumar 0003, Ke Tan 0001, Buye Xu, DeLiang Wang |
INTERSPEECH | 5 |
| 2023 | Multi-input Multi-output Complex Spectral Mapping for Speaker Separation
Hassan Taherian, Ashutosh Pandey 0004, Buye Xu, DeLiang Wang |
INTERSPEECH | 5 |
| 2023 | Time-Domain Speech Enhancement for Robust Automatic Speech RecognitionabstractIt has been shown that the intelligibility of noisy speech can be improved by speech enhancement algorithms. However, speech enhancement has not been established as an effective frontend for robust automatic speech recognition (ASR) in noisy conditions compared to an ASR model trained on noisy speech directly. The divide between speech enhancement and ASR impedes the progress of robust ASR systems especially as speech enhancement has made big strides in recent years. In this work, we focus on eliminating this divide with an ARN (attentive recurrent network) based time-domain enhancement model. The proposed system fully decouples speech enhancement and an acoustic model trained only on clean speech. Results on the CHiME-2 corpus show that ARN enhanced speech translates to improved ASR results. The proposed system achieves 6.28% average word error rate, outperforming the previous best by 19.3% relatively. Ashutosh Pandey 0004, DeLiang Wang |
INTERSPEECH | 3 |
| 2023 | Announcement of the Neural Networks Best Paper Award
Taro Toyoizumi, DeLiang Wang |
Neural Networks | 2 |
| 2023 | Another bumper year
Taro Toyoizumi, DeLiang Wang |
Neural Networks | 2 |
| 2023 | Deep MCANC: A deep learning approach to multi-channel active noise control
Hao Zhang 0112, DeLiang Wang |
Neural Networks | 2 |
| 2023 | Attentive Training: A New Training Framework for Speech EnhancementabstractDealing with speech interference in a speech enhancement system requires either speaker separation or target speaker extraction. Speaker separation has multiple output streams with arbitrary assignments while target speaker extraction requires additional cueing for speaker selection. Both of these are not suitable for a standalone speech enhancement system with one output stream. In this study, we propose a novel training framework, calledAttentive Training, to extend speech enhancement to deal with speech interruptions. Attentive training is based on the observation that, in the real world, multiple talkers very unlikely start speaking at the same time, and therefore, a deep neural network can be trained to create a representation of the first speaker and utilize it to attend to or track that speaker in a multitalker noisy mixture. We present experimental results and comparisons to demonstrate the effectiveness of attentive training for speech enhancement. Ashutosh Pandey 0004, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Low-Latency Active Noise Control Using Attentive Recurrent NetworkabstractProcessing latency is a critical issue for active noise control (ANC) due to the causality constraint of ANC systems. This paper addresses low-latency ANC in the context of deep learning (i.e. deep ANC). A time-domain method using an attentive recurrent network (ARN) is employed to perform deep ANC with smaller frame sizes, thus reducing algorithmic latency of deep ANC. In addition, we introduce a delay-compensated training to perform ANC using predicted noise for several milliseconds. Moreover, a revised overlap-add method is utilized during signal resynthesis to avoid the latency introduced due to overlaps between neighboring time frames. Experimental results show the effectiveness of the proposed strategies for achieving low-latency deep ANC. Combining the proposed strategies is capable of yielding zero, even negative, algorithmic latency without affecting ANC performance much, thus alleviating the causality constraint in ANC design. Hao Zhang 0112, Ashutosh Pandey 0004, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | $F0$ Estimation and Voicing Detection With Cascade Architecture in Noisy SpeechabstractAs a fundamental problem in speech processing, pitch tracking has been studied for decades. While strong performance has been achieved on clean speech, pitch tracking in noisy speech is still challenging. Severe non-stationary noises not only corrupt the harmonic structure in voiced intervals but also make it difficult to determine the existence of voiced speech. Given the importance of voicing detection for pitch tracking, this study proposes a neural cascade architecture that jointly performs pitch estimation and voicing detection. The cascade architecture optimizes a speech enhancement module and a pitch tracking module, and is trained in a speaker-independent and noise-independent way. It is observed that incorporating the enhancement module improves both pitch estimation and voicing detection accuracy, especially in low signal-to-noise ratio (SNR) conditions. In addition, compared with frameworks that combine corresponding single-task models, the proposed multi-task framework achieves better performance and is more efficient. Experimental results show that the proposed method is robust to different noise conditions and substantially outperforms other competitive pitch tracking methods. Yixuan Zhang 0005, Heming Wang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech EnhancementabstractIn this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes speech signals from all channels independently using a dual-path attentive recurrent network (ARN), which is a recurrent neural network (RNN) augmented with self-attention. Next, an ARN is introduced along the spatial dimension for spatial context aggregation. TPARN is designed as a multiple-input and multiple-output architecture to enhance all input channels simultaneously. Experimental results demonstrate the superiority of TPARN over existing state-of-the-art approaches. Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang |
ICASSP | 6 |
| 2022 | Multichannel Speech Enhancement Without BeamformingabstractDeep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post-filter in a multistage processing provides additional improvements. In this work, we propose a two-stage strategy for multi-channel speech enhancement that does not require a traditional beamformer for additional performance. First, we propose a novel attentive dense convolutional network (ADCN) for estimating real and imaginary parts of complex spectrogram. ADCN obtains state-of-the-art results among single-stage models. Next, we use ADCN with a recently proposed triple-path attentive recurrent network (TPARN) for estimating waveform samples. The proposed strategy uses two insights; first, using different approaches in two stages; and second, using a stronger model in the first stage. We illustrate the efficacy of our strategy by evaluating multiple models in a two-stage approach with and without a traditional beamformer. Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang |
ICASSP | 6 |
| 2022 | Location-Based Training for Multi-Channel Talker-Independent Speaker SeparationabstractPermutation-invariant training (PIT) is a dominant approach for addressing the permutation ambiguity problem in talker-independent speaker separation. Leveraging spatial information afforded by microphone arrays, we propose a new training approach to resolving permutation ambiguities for multi-channel speaker separation. The proposed approach, named location-based training (LBT), assigns speakers on the basis of their spatial locations. This training strategy is easy to apply, and organizes speakers according to their positions in physical space. Specifically, this study investigates azimuth angles and source distances for location-based training. Evaluation results on separating two- and three-speaker mixtures show that azimuth-based training consistently outperforms PIT, and distance-based training further improves the separation performance when speaker azimuths are close. Furthermore, we dynamically select azimuth-based or distance-based training by estimating the azimuths of separated speakers, which further improves separation performance. LBT has a linear training complexity with respect to the number of speakers, as opposed to the factorial complexity of PIT. We further demonstrate the effectiveness of LBT for the separation of four and five concurrent speakers. Hassan Taherian, Ke Tan 0001, DeLiang Wang |
ICASSP | 3 |
| 2022 | Improving Noise Robustness of Contrastive Speech Representation Learning with Speech ReconstructionabstractNoise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this work, instead of suppressing background noise with a conventional cascaded pipeline, we employ a noise-robust representation learned by a refined self-supervised framework for noisy speech recognition. We propose to combine a reconstruction module with contrastive learning and perform multi-task continual pre-training on noisy data. The reconstruction module is used for auxiliary learning to improve the noise robustness of the learned representation and thus is not required during inference. Experiments demonstrate the effectiveness of our proposed method. Our model substantially reduces the word error rate (WER) for the synthesized noisy LibriSpeech test sets, and yields around 4.1/7.5% WER reduction on noisy clean/other test sets compared to data augmentation. For the real-world noisy speech from the CHiME-4 challenge (1-channel track), we have obtained the state of the art ASR performance without any denoising front-end. Moreover, we achieve comparable performance to the best supervised approach reported with only 16% of labeled data. Heming Wang, Yao Qian, Xiaofei Wang 0009, Chengyi Wang 0002, Shujie Liu 0001, Takuya Yoshioka, Jinyu Li 0001, DeLiang Wang |
ICASSP | 9 |
| 2022 | Localization based Sequential Grouping for Continuous Speech SeparationabstractThis study investigates robust speaker localization for continuous speech separation and speaker diarization, where we use speaker directions to group non-contiguous segments of the same speaker. Assuming that speakers do not move and are located in different directions, the direction of arrival (DOA) information provides an informative cue for accurate sequential grouping and speaker diarization. Our system is block-online in the following sense. Given a block of frames with at most two speakers, we apply a two-speaker separation model to separate (and enhance) the speakers, estimate the DOA of each separated speaker, and group the separation results across blocks based on the DOA estimates. Speaker diarization and speaker-attributed speech recognition results on the LibriCSS corpus demonstrate the effectiveness of the proposed algorithm. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2022 | Cross-Domain Speech Enhancement with a Neural Cascade ArchitectureabstractThis paper proposes a novel cascade architecture to address the monaural speech enhancement problem. We leverage three different domains of speech representation, namely spectral magnitude, waveform, and complex spectrogram, to progressively suppress the background noise within noisy speech. Our proposed neural cascade architecture consists of three modules, and each operates on the original noisy input and the output of the previous module in a distinct speech representation. During training, the network simultaneously optimizes all modules with a triple-domain loss. Experiments on the WSJ0 SI-84 corpus demonstrate that our proposed approach achieves superior enhancement results, and substantially outperforms previous baselines in terms of both speech quality and intelligibility. Heming Wang, DeLiang Wang |
ICASSP | 2 |
| 2022 | Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex DomainabstractBone-conduction (BC) microphones capture speech signals by converting the vibrations of the human skull into electrical signals. BC sensors are insensitive to acoustic noise, but limited in bandwidth. On the other hand, conventional or air-conduction (AC) microphones are capable of capturing full-band speech, but are susceptible to background noise. We propose to combine the strengths of AC and BC microphones by employing a convolutional recurrent network that performs complex spectral mapping. To better utilize signals from both kinds of microphone, we employ attention-based fusion with early-fusion and late-fusion strategies. Experiments demonstrate the superiority of the proposed method over other recent speech enhancement methods combining BC and AC signals. In addition, our enhancement performance is significantly better than conventional speech enhancement counterparts, especially in low signal-to-noise ratio scenarios. Heming Wang, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 3 |
| 2022 | Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand ChallengeabstractThe ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions. Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu |
ICASSP | 10 |
| 2022 | Neural Cascade Architecture for Joint Acoustic Echo and Noise SuppressionabstractIn this paper, we propose a neural cascade architecture for joint acoustic echo and noise suppression. The proposed cascade architecture consists of two modules. A convolutional recurrent network (CRN) is employed in the first module for complex spectral mapping. The output is then fed as an additional input to the second module, where a long short-term memory network (LSTM) is utilized for magnitude mask estimation. The entire architecture is trained in an end-to-end manner with the two modules optimized jointly using a single loss function. The final output is generated using the enhanced phase and magnitude obtained from the first and the second module, respectively. The cascade architecture enables the proposed method to obtain robust magnitude estimation as well as phase enhancement. Evaluation results show that the proposed method effectively suppresses acoustic echo and noise while preserving good speech quality, and significantly outperforms related methods. Hao Zhang 0112, DeLiang Wang |
ICASSP | 2 |
| 2022 | Attentive Training: A New Training Framework for Talker-independent Speaker Extraction
Ashutosh Pandey 0004, DeLiang Wang |
INTERSPEECH | 2 |
| 2022 | Time-domain Ad-hoc Array Speech Enhancement Using a Triple-path NetworkabstractDeep neural networks (DNNs) are very effective for multichannel speech enhancement with fixed array geometries.However, it is not trivial to use DNNs for ad-hoc arrays with unknown order and placement of microphones.We propose a novel triplepath network for ad-hoc array processing in the time domain.The key idea in the network design is to divide the overall processing into spatial processing and temporal processing and use self-attention for spatial processing.Using self-attention for spatial processing makes the network invariant to the order and the number of microphones.The temporal processing is done independently for all channels using a recently proposed dual-path attentive recurrent network.The proposed network is a multiple-input multiple-output architecture that can simultaneously enhance signals at all microphones.Experimental results demonstrate the excellent performance of the proposed approach.Further, we present analysis to demonstrate the effectiveness of the proposed network in utilizing multichannel information even from microphones at far locations. Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang |
INTERSPEECH | 6 |
| 2022 | Neural Vocoder is All You Need for Speech Super-resolutionabstractSpeech super-resolution (SR) is a task to increase speech sampling rate by generating high-frequency components. Existing speech SR methods are trained in constrained experimental settings, such as a fixed upsampling ratio. These strong constraints can potentially lead to poor generalization ability in mismatched real-world cases. In this paper, we propose a neural vocoder based speech super-resolution method (NVSR) that can handle a variety of input resolution and upsampling ratios. NVSR consists of a mel-bandwidth extension module, a neural vocoder module, and a post-processing module. Our proposed system achieves state-of-the-art results on the VCTK multi-speaker benchmark. On 44.1 kHz target resolution, NVSR outperforms WSRGlow and Nu-wave by 8% and 37% respectively on log spectral distance and achieves a significantly better perceptual quality. We also demonstrate that prior knowledge in the pre-trained vocoder is crucial for speech SR by performing mel-bandwidth extension with a simple replication-padding method. Samples can be found in https://haoheliu.github.io/nvsr. Haohe Liu, Woosung Choi, Xubo Liu 0001, Qiuqiang Kong, Qiao Tian 0001, DeLiang Wang |
INTERSPEECH | 6 |
| 2022 | VoiceFixer: A Unified Framework for High-Fidelity Speech RestorationabstractSpeech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, speech signals can be degraded by several different distortions simultaneously in the real world. It is thus important to extend speech restoration models to deal with multiple distortions. In this paper, we introduce VoiceFixer, a unified framework for high-fidelity speech restoration. VoiceFixer restores speech from multiple distortions (e.g., noise, reverberation, and clipping) and can expand degraded speech (e.g., noisy speech) with a low bandwidth to 44.1 kHz full-bandwidth high-fidelity speech. We design VoiceFixer based on (1) an analysis stage that predicts intermediate-level features from the degraded speech, and (2) a synthesis stage that generates waveform using a neural vocoder. Both objective and subjective evaluations show that VoiceFixer is effective on severely degraded speech, such as real-world historical speech recordings. Samples of VoiceFixer are available at https://haoheliu.github.io/voicefixer. Haohe Liu, Xubo Liu 0001, Qiuqiang Kong, Qiao Tian 0001, Yan Zhao 0010, DeLiang Wang, Chuanzeng Huang, Yuxuan Wang 0002 |
INTERSPEECH | 6 |
| 2022 | Attentive Recurrent Network for Low-Latency Active Noise ControlabstractProcessing latency is a critical issue for active noise control (ANC) due to the causality constraint of ANC systems. This paper addresses low-latency ANC in the deep learning framework (i.e. deep ANC). A time-domain method using an attentive recurrent network is employed to perform deep ANC with smaller frame sizes, thus reducing algorithmic latency of deep ANC. In addition, a delay-compensated training strategy is introduced to perform ANC using predicted noise for several milliseconds. Moreover, we utilize a revised overlap-add method during signal resynthesis to avoid the latency introduced due to overlaps between neighboring time frames. Experimental results show that the proposed strategies are effective for achieving low-latency deep ANC. Combining the proposed strategies is capable of yielding zero, even negative, algorithmic latency without significantly affecting ANC performance. Hao Zhang 0112, Ashutosh Pandey 0004, DeLiang Wang |
INTERSPEECH | 3 |
| 2022 | Densely-connected Convolutional Recurrent Network for Fundamental Frequency Estimation in Noisy Speechabstractin turn. Experimental results show that the cascade model brings further improvements to the DC-CRN model, especially in low signal-to-noise ratio (SNR) conditions. Yixuan Zhang 0005, Heming Wang, DeLiang Wang |
INTERSPEECH | 3 |
| 2022 | Continual growth and a transition
Kenji Doya, Taro Toyoizumi, DeLiang Wang |
Neural Networks | 3 |
| 2022 | Announcement of the Neural Networks Best Paper Award
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2022 | Self-Attending RNN for Speech Enhancement to Improve Cross-Corpus GeneralizationabstractDeep neural networks (DNNs) represent the mainstream methodology for supervised speech enhancement, primarily due to their capability to model complex functions using hierarchical representations. However, a recent study revealed that DNNs trained on a single corpus fail to generalize to untrained corpora, especially in low signal-to-noise ratio (SNR) conditions. Developing a noise, speaker, and corpus independent speech enhancement algorithm is essential for real-world applications. In this study, we propose a self-attending recurrent neural network (SARNN) for time-domain speech enhancement to improve cross-corpus generalization. SARNN comprises of recurrent neural networks (RNNs) augmented with self-attention blocks and feedforward blocks. We evaluate SARNN on different corpora with nonstationary noises in low SNR conditions. Experimental results demonstrate that SARNN substantially outperforms competitive approaches to time-domain speech enhancement, such as RNNs and dual-path SARNNs. Additionally, we report an important finding that the two popular approaches to speech enhancement: complex spectral mapping and time-domain enhancement, obtain similar results for RNN and SARNN with large-scale training. We also provide a challenging subset of the test set used in this study for evaluating future algorithms and facilitating direct comparisons. Ashutosh Pandey 0004, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Multi-Channel Talker-Independent Speaker Separation Through Location-Based TrainingabstractPermutation ambiguity is a crucial issue for deep learning based talker-independent speaker separation. Deep clustering and permutation invariant training (PIT) have been widely used to address the permutation ambiguity problem in monaural scenarios. Although both approaches have been extended to multi-microphone scenarios, we believe that the permutation ambiguity problem can be naturally avoided by leveraging the spatial relations of multiple speakers. In this study, we present location-based training (LBT), a new approach to achieve talker independency in multi-channel speaker separation. Unlike PIT that examines all possible permutations, LBT assigns speakers according to their positions in physical space. With a linear training complexity to the number of concurrent speakers, LBT is computationally much more efficient than PIT with a factorial complexity, particularly when a large number of overlapping speakers needs to be separated. Specifically, we propose two training criteria: azimuth-based and distance-based training, using speaker azimuths and distances relative to a microphone array, respectively. Evaluation results show that LBT significantly outperforms PIT on two-speaker and three-speaker mixtures with different array geometries and in various acoustic conditions. In addition, we propose a joint training strategy to integrate azimuth-based and distance-based training, which further improves separation performance. Hassan Taherian, Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Neural Spectrospatial FilteringabstractAs the most widely-used spatial filtering approach for multi-channel speech separation, beamforming extracts the target speech signal arriving from a specific direction. An emerging alternative approach is multi-channel complex spectral mapping, which trains a deep neural network (DNN) to directly estimate the real and imaginary spectrograms of the target speech signal from those of the multi-channel noisy mixture. In this all-neural approach, the trained DNN itself becomes a nonlinear, time-varying spectrospatial filter. However, it remains unclear how this approach performs relative to commonly-used beamforming techniques on different array configurations and acoustic environments. This paper is devoted to examining this issue in a systematic way. Comprehensive evaluations show that multi-channel complex spectral mapping achieves separation performance comparable to or better than beamforming for different array geometries and speech separation tasks and reduces to monaural complex spectral mapping in single-channel conditions, demonstrating the general utility of this approach on multi-channel and single-channel speech separation. In addition, such an approach is computationally more efficient than widely-used mask-based beamforming. We conclude that this neural spectrospatial filter provides a strong alternative to traditional and mask-based beamforming. Ke Tan 0001, Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Neural Cascade Architecture With Triple-Domain Loss for Speech EnhancementabstractThis paper proposes a neural cascade architecture to address the monaural speech enhancement problem. The cascade architecture is composed of three modules which optimize in turn enhanced speech with respect to the magnitude spectrogram, the time-domain signal and the complex spectrogram. Each module takes as input the noisy speech and the output obtained from the previous module, and generates a prediction of the respective target. Our model is trained in an end-to-end manner, using a triple-domain loss function that accounts for three domains of signal representation. Experimental results on the WSJ0 SI-84 corpus show that the proposed model outperforms other strong speech enhancement baselines in terms of objective speech quality and intelligibility. Heming Wang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Fusing Bone-Conduction and Air-Conduction Sensors for Complex-Domain Speech EnhancementabstractSpeech enhancement aims to improve the listening quality and intelligibility of noisy speech in adverse environments. It proves to be challenging to perform speech enhancement in very low signal-to-noise ratio (SNR) conditions. Conventional speech enhancement utilizes air-conduction (AC) microphones, which are sensitive to background noise but capable of capturing full-band signals. On the other hand, bone-conduction (BC) sensors are unaffected by acoustic noise, but recorded speech has limited bandwidth. This study proposes an attention-based fusion method to combine the strengths of AC and BC signals and perform complex spectral mapping for speech enhancement. Experiments on the EMSB dataset demonstrate that the proposed approach effectively leverages the advantages of AC and BC sensors, and outperforms a recent time-domain baseline in all conditions. We also show that the sensor fusion method is superior to single-sensor counterparts, especially in low SNR conditions. As the amount of BC data is very limited, we additionally propose a semi-supervised technique to utilize both parallelly and unparallely recorded AC and BC speech signals. With additional AC speech from the AISHELL-1 dataset, we achieve similar performance to supervised learning with only 50% parallel data. Heming Wang, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Neural Cascade Architecture for Multi-Channel Acoustic Echo SuppressionabstractTraditional acoustic echo cancellation (AEC) works by identifying an acoustic impulse response using adaptive algorithms. This paper proposes a neural cascade architecture for joint acoustic echo and noise suppression to address both single-channel and multi-channel AEC (MCAEC) problems. The proposed cascade architecture consists of two modules. A convolutional recurrent network (CRN) is employed in the first module for complex spectral mapping. Its output is fed as an additional input to the second module, where a long short-term memory network (LSTM) is utilized for magnitude mask estimation. The entire architecture is trained in an end-to-end manner with the two modules optimized jointly using a single loss function. The final output is generated using the enhanced phase and magnitude obtained from the first and the second module, respectively. The cascade architecture enables the proposed method to obtain robust magnitude estimation as well as phase enhancement. The proposed method is investigated under different AEC setups. We find that the deep learning based approach avoids the no-uniqueness problem in traditional MCAEC. For MCAEC setups with multiple microphones, combining deep MCAEC with supervised beamforming further improves the system performance. Evaluation results show that the proposed approach effectively suppresses acoustic echo and noise while preserving speech quality, and consistently outperforms related methods under different setups. Hao Zhang 0112, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Time-Domain Loss Modulation Based on Overlap Ratio for Monaural Conversational Speaker SeparationabstractExisting speaker separation methods deliver excellent performance on fully overlapped signal mixtures. To apply these methods in daily conversations that include occasional concurrent speakers, recent studies incorporate both overlapped and non-overlapped segments in the training data. However, such training data can degrade the separation performance due to triviality of non-overlapped segments where the model reflects the input to the output. We propose a new loss function for speaker separation based on permutation invariant training that dynamically reweighs losses using the segment overlap ratio. The new loss function emphasizes overlapped regions while deemphasizing the segments with single speakers. We demonstrate the effectiveness of the proposed loss function on an automatic speech recognition (ASR) task. Experiments on the recently introduced LibriCSS corpus show that our proposed single-channel method produces consistent improvements compared to baseline methods. Hassan Taherian, DeLiang Wang |
ICASSP | 2 |
| 2021 | Compressing Deep Neural Networks for Efficient Speech EnhancementabstractThe use of deep neural networks (DNNs) has dramatically improved the performance of speech enhancement in the past decade. However, a large DNN is typically required to achieve strong enhancement performance, and this kind of model is both computationally intensive and memory consuming. Hence it is difficult to deploy such DNNs on devices with limited hardware resources or in applications with strict latency requirements. In order to address this problem, we propose a model compression pipeline to reduce DNN size for speech enhancement, which is based on three kinds of techniques: sparse regularization, iterative pruning and clustering-based quantization. Evaluation results show that our approach substantially reduces the sizes of different DNNs without significantly affecting their enhancement performance. Moreover, we find that training and compressing a large DNN yields higher STOI and PESQ than directly training a small DNN that has a comparable size to the compressed DNN. This further suggests the benefits of using the proposed model compression approach. Ke Tan 0001, DeLiang Wang |
ICASSP | 2 |
| 2021 | Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral MappingabstractSpeech quality and intelligibility can be severely degraded by back-ground noise in mobile communication. In order to attenuate back-ground noise, speech enhancement systems have been integrated into mobile phones, and a microphone array is typically deployed to improve the enhancement performance. This paper proposes a novel approach to real-time speech enhancement for dual-microphone mobile phones. Our approach employs a causal densely-connected convolutional recurrent network to perform dual-channel complex spectral mapping. We apply a structured pruning technique for compressing the model without significantly affecting the enhancement performance. This leads to a real-time enhancement system for on-device processing. Evaluation results show that the pro-posed approach substantially advances the performance of an earlier approach to dual-channel speech enhancement for mobile communication. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 3 |
| 2021 | Count And Separate: Incorporating Speaker Counting For Continuous Speaker SeparationabstractThis study leverages frame-wise speaker counting to switch between speech enhancement and speaker separation for continuous speaker separation. The proposed approach counts the number of speakers at each frame. If there is no speaker overlap, a speech enhancement model is used to suppress noise and reverberation. Otherwise, a speaker separation model based on permutation invariant training is utilized to separate multiple speakers in noisy-reverberant conditions. We stitch the results from the enhancement and separation models based on their predictions in a small augmented window of frames surrounding an overlapped segment. Assuming a fixed array geometry between training and testing, we use multi-microphone complex spectral mapping for enhancement and separation, where deep neural networks are trained to predict the real and imaginary (RI) components of direct sound from stacked reverberant-noisy RI components of multiple microphones. Experimental results on the LibriCSS dataset demonstrate the effectiveness of our approach. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2021 | Complex Ratio Masking For Singing Voice SeparationabstractMusic source separation is important for applications such as karaoke and remixing. Much of previous research focuses on estimating short-time Fourier transform (STFT) magnitude and discarding phase information. We observe that, for singing voice separation, phase can make considerable improvement in separation quality. This paper proposes a complex ratio masking method for voice and accompaniment separation. The proposed method employs DenseUNet with self attention to estimate the real and imaginary components of STFT for each sound source. A simple ensemble technique is introduced to further improve separation performance. Evaluation results demonstrate that the proposed method outperforms recent state-of-the-art models for both separated voice and accompaniment. Yixuan Zhang 0005, DeLiang Wang |
ICASSP | 3 |
| 2021 | A Deep Learning Method to Multi-Channel Active Noise Control
Hao Zhang 0112, DeLiang Wang |
Interspeech | 2 |
| 2021 | A Deep Learning Approach to Multi-Channel and Multi-Microphone Acoustic Echo Cancellation
Hao Zhang 0112, DeLiang Wang |
Interspeech | 2 |
| 2021 | Maintaining the Publication Infrastructure in a Worldwide Pandemic
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2021 | Deep ANC: A deep learning approach to active noise control
Hao Zhang 0112, DeLiang Wang |
Neural Networks | 2 |
| 2021 | Recurrent Neural Networks and Acoustic Features for Frame-Level Signal-to-Noise Ratio EstimationabstractIt is important to know the presence and the relative level of background noise for many speech processing tasks. Frame-level signal-to-noise ratio (SNR) provides a measure of instantaneous noise level of a noisy signal, and its estimation has been researched for decades. This problem can be approached from a supervised learning perspective by predicting SNR from features of noisy speech. In this study, we introduce a deep learning algorithm for frame-level SNR estimation. The proposed algorithm employs recurrent neural networks (RNNs) with long short-term memory (LSTM) to leverage contextual information. We also systematically examine a range of acoustic features and investigate feature combinations using Group Lasso and sequential floating forward selection (SFFS). The proposed algorithm naturally leads to an utterance-level SNR estimator. Systematical evaluations show that the proposed algorithm provides an accurate estimate of frame-level SNR, as well as utterance-level SNR, under different noise conditions, outperforming other estimators. Hao Li 0046, DeLiang Wang, Xueliang Zhang 0001, Guanglai Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Dense CNN With Self-Attention for Time-Domain Speech EnhancementabstractSpeech enhancement in the time domain is becoming increasingly popular in recent years, due to its capability to jointly enhance both the magnitude and the phase of speech. In this work, we propose a dense convolutional network (DCN) with self-attention for speech enhancement in the time domain. DCN is an encoder and decoder based architecture with skip connections. Each layer in the encoder and the decoder comprises a dense block and an attention module. Dense blocks and attention modules help in feature extraction using a combination of feature reuse, increased network depth, and maximum context aggregation. Furthermore, we reveal previously unknown problems with a loss based on the spectral magnitude of enhanced speech. To alleviate these problems, we propose a novel loss based on magnitudes of enhanced speech and a predicted noise. Even though the proposed loss is based on magnitudes only, a constraint imposed by noise prediction ensures that the loss enhances both magnitude and phase. Experimental results demonstrate that DCN trained with the proposed loss substantially outperforms other state-of-the-art approaches to causal and non-causal speech enhancement. Ashutosh Pandey 0004, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Towards Model Compression for Deep Learning Based Speech EnhancementabstractThe use of deep neural networks (DNNs) has dramatically elevated the performance of speech enhancement over the last decade. However, to achieve strong enhancement performance typically requires a large DNN, which is both memory and computation consuming, making it difficult to deploy such speech enhancement systems on devices with limited hardware resources or in applications with strict latency requirements. In this study, we propose two compression pipelines to reduce the model size for DNN-based speech enhancement, which incorporates three different techniques: sparse regularization, iterative pruning and clustering-based quantization. We systematically investigate these techniques and evaluate the proposed compression pipelines. Experimental results demonstrate that our approach reduces the sizes of four different models by large margins without significantly sacrificing their enhancement performance. In addition, we find that the proposed approach performs well on speaker separation, which further demonstrates the effectiveness of the approach for compressing speech separation models. Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Deep Learning Based Real-Time Speech Enhancement for Dual-Microphone Mobile PhonesabstractIn mobile speech communication, speech signals can be severely corrupted by background noise when the far-end talker is in a noisy acoustic environment. To suppress background noise, speech enhancement systems are typically integrated into mobile phones, in which one or more microphones are deployed. In this study, we propose a novel deep learning based approach to real-time speech enhancement for dual-microphone mobile phones. The proposed approach employs a new densely-connected convolutional recurrent network to perform dual-channel complex spectral mapping. We utilize a structured pruning technique to compress the model without significantly degrading the enhancement performance, which yields a low-latency and memory-efficient enhancement system for real-time processing. Experimental results suggest that the proposed approach consistently outperforms an earlier approach to dual-channel speech enhancement for mobile phone communication, as well as a deep learning based beamformer. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Speaker Separation Using Speaker Inventories and Estimated SpeechabstractWe propose speaker separation using speaker inventories and estimated speech (SSUSIES), a framework leveraging speaker profiles and estimated speech for speaker separation. SSUSIES contains two methods, speaker separation using speaker inventories (SSUSI) and speaker separation using estimated speech (SSUES). SSUSI performs speaker separation with the help of speaker inventory. By combining the advantages of permutation invariant training (PIT) and speech extraction, SSUSI significantly outperforms conventional approaches. SSUES is a widely applicable technique that can substantially improve speaker separation performance using the output of first-pass separation. We evaluate the models on both speaker separation and speech recognition metrics. Zhuo Chen 0006, DeLiang Wang, Jinyu Li 0001, Yifan Gong 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Towards Robust Speech Super-ResolutionabstractSpeech super-resolution (SR) aims to increase the sampling rate of a given speech signal by generating high-frequency components. This paper proposes a convolutional neural network (CNN) based SR model that takes advantage of information from both time and frequency domains. Specifically, the proposed CNN is a time-domain model that takes the raw waveform of low-resolution speech as the input, and outputs an estimate of the corresponding high-resolution waveform. During the training stage, we employ a cross-domain loss to optimize the network. We compare our model with several deep neural network (DNN) based SR models, and experiments show that our model outperforms existing models. Furthermore, the robustness of DNN-based models is investigated, in particular regarding microphone channels and downsampling schemes, which have a major impact on the performance of DNN-based SR models. By training with proper datasets and preprocessing, we improve the generalization capability for untrained microphone channels and unknown downsampling schemes. Heming Wang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Multi-microphone Complex Spectral Mapping for Utterance-wise and Continuous Speech SeparationabstractWe propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation and dereverberation. Our study first investigates offline utterance-wise speaker separation and then extends to block-online continuous speech separation (CSS). Assuming a fixed array geometry between training and testing, we train deep neural networks (DNN) to predict the real and imaginary (RI) components of target speech at a reference microphone from the RI components of multiple microphones. We then integrate multi-microphone complex spectral mapping with minimum variance distortionless response (MVDR) beamforming and post-filtering to further improve separation, and combine it with frame-level speaker counting for block-online CSS. Although our system is trained on simulated room impulse responses (RIR) based on a fixed number of microphones arranged in a given geometry, it generalizes well to a real array with the same geometry. State-of-the-art separation performance is obtained on the simulated two-talker SMS-WSJ corpus and the real-recorded LibriCSS dataset. Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Densely Connected Neural Network with Dilated Convolutions for Real-Time Speech Enhancement in The Time DomainabstractIn this work, we propose a fully convolutional neural network for real-time speech enhancement in the time domain. The proposed network is an encoder-decoder based architecture with skip connections. The layers in the encoder and the decoder are followed by densely connected blocks comprising of dilated and causal convolutions. The dilated convolutions help in context aggregation at different resolutions. The causal convolutions are used to avoid information flow from future frames, hence making the network suitable for real-time applications. We also propose to use sub-pixel convolutional layers in the decoder for upsampling. Further, the model is trained using a loss function with two components; a time-domain loss and a frequency-domain loss. The proposed loss function outperforms the time-domain loss. Experimental results show that the proposed model significantly outperforms other real-time state-of-the-art models in terms of objective intelligibility and quality scores. Ashutosh Pandey 0004, DeLiang Wang |
ICASSP | 2 |
| 2020 | Talker-Independent Speaker Separation in Reverberant ConditionsabstractSpeaker separation refers to the task of separating a mixture signal comprising two or more speakers. Impressive advances have been made recently in deep learning based talker-independent speaker separation. But such advances are achieved in anechoic conditions. We address talker-independent speaker separation in reverberant conditions by exploring a recently proposed deep CASA approach. To effectively deal with speaker separation and speech dereverberation, we propose a two-stage strategy where reverberant utterances are first separated and then dereverberated. The two-stage deep CASA method outperforms other talker-independent separation methods. In addition, the deep CASA algorithm produces substantial speech intelligibility improvements for human listeners, with a particularly large benefit for hearing-impaired listeners. Masood Delfarah, DeLiang Wang |
ICASSP | 3 |
| 2020 | Deep Casa for Talker-independent Monaural Speech SeparationabstractMonaural speech separation is the task of separating target speech from interference in single-channel recordings. Although substantial progress has been made recently in deep learning based speech separation, previous studies usually focus on a single type of interference, either background noise or competing speakers. In this study, we address both speech and nonspeech interference, i.e., monaural speaker separation in noise, in a talker-independent fashion. We extend a recently proposed deep CASA system to deal with noisy speaker mixtures. To facilitate speech enhancement, a denoising module is added to deep CASA as a front-end processor. The proposed systems achieve state-of-the-art results on a benchmark noisy two-speaker separation dataset. The denoising module leads to substantial performance gain across various noise types, and even better generalization in noise-free conditions. Masood Delfarah, DeLiang Wang |
ICASSP | 3 |
| 2020 | Improving Robustness of Deep Learning Based Monaural Speech Enhancement Against Processing ArtifactsabstractIn voice telecommunication, the intelligibility and quality of speech signals can be severely degraded by background noise if the speaker at the transmitting end talks in a noisy environment. Therefore, a speech enhancement system is typically integrated into the transmitter device or the receiver device. Without the knowledge of whether the other end is equipped with a speech enhancer, the transmitter and receiver devices can both process a speech signal with their speech enhancers. In this study, we find that enhancing a speech signal twice can dramatically degrade the enhancement performance. This is because the downstream speech enhancer is sensitive to the processing artifacts introduced by the upstream enhancer. We analyze this problem and propose a new training scheme for the downstream deep learning based speech enhancement model. Our experimental results show that the proposed training strategy substantially elevate the robustness of speech enhancers against artifacts induced by another speech enhancer. Ke Tan 0001, DeLiang Wang |
ICASSP | 2 |
| 2020 | Multi-Microphone Complex Spectral Mapping for Speech DereverberationabstractThis study proposes a multi-microphone complex spectral mapping approach for speech dereverberation on a fixed array geometry. In the proposed approach, a deep neural network (DNN) is trained to predict the real and imaginary (RI) components of direct sound from the stacked reverberant (and noisy) RI components of multiple microphones. We also investigate the integration of multi-microphone complex spectral mapping with beamforming and post-filtering. Experimental results on multi-channel speech dereverberation demonstrate the effectiveness of the proposed approach. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2020 | Time-Frequency Loss for CNN Based Speech Super-ResolutionabstractSpeech super-resolution (SR), also called speech bandwidth extension (BWE), aims to increase the sampling rate of a given lower resolution speech signal. Recent years have witnessed the successful application of deep neural networks in time or frequency domains, and deep learning has improved the performance considerably compared with conventional approaches. This paper proposes an autoencoder based fully convolutional neural network (CNN) that merges the information from both time and frequency domains. At the training time, we optimize the CNN using a new time-frequency loss (T-F loss), which combines a time domain loss and a frequency domain loss. The experimental results show that our model trained with the T-F loss achieves significantly better results than other state-of-the-art models, and yields balanced performance in terms of time and frequency metrics. Heming Wang, DeLiang Wang |
ICASSP | 2 |
| 2020 | Learning Complex Spectral Mapping for Speech Enhancement with Improved Cross-Corpus Generalization
Ashutosh Pandey 0004, DeLiang Wang |
INTERSPEECH | 2 |
| 2020 | Noisy-Reverberant Speech Enhancement Using DenseUNet with Time-Frequency Attention
Yan Zhao 0010, DeLiang Wang |
INTERSPEECH | 2 |
| 2020 | Frame-Level Signal-to-Noise Ratio Estimation Using Deep Learning
Hao Li 0046, DeLiang Wang, Xueliang Zhang 0001, Guanglai Gao |
INTERSPEECH | 2 |
| 2020 | A Deep Learning Approach to Active Noise Control
Hao Zhang 0112, DeLiang Wang |
INTERSPEECH | 2 |
| 2020 | Announcement of the Neural Networks Best Paper Award
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2020 | Causal Deep CASA for Monaural Talker-Independent Speaker SeparationabstractTalker-independent monaural speaker separation aims to separate concurrent speakers from a single-microphone recording. Inspired by human auditory scene analysis (ASA) mechanisms, a two-stage deep CASA approach has been proposed recently to address this problem, which achieves state-of-the-art results in separating mixtures of two or three speakers. A main limitation of deep CASA is that it is a non-causal system, while many speech processing applications, e.g., telecommunication and hearing prosthesis, require causal processing. In this study, we propose a causal version of deep CASA to address this limitation. First, we modify temporal connections, normalization and clustering algorithms in deep CASA so that no future information is used throughout the deep network. We then train a C-speaker (C ≥ 2) deep CASA system in a speaker-number-independent fashion, generalizable to speech mixtures with up to C speakers without the prior knowledge about the speaker number. Experimental results show that causal deep CASA achieves excellent speaker separation performance with known or unknown speaker numbers. DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | On Cross-Corpus Generalization of Deep Learning Based Speech EnhancementabstractIn recent years, supervised approaches using deep neural networks (DNNs) have become the mainstream for speech enhancement. It has been established that DNNs generalize well to untrained noises and speakers if trained using a large number of noises and speakers. However, we find that DNNs fail to generalize to new speech corpora in low signal-to-noise ratio (SNR) conditions. In this work, we establish that the lack of generalization is mainly due to the channel mismatch, i.e. different recording conditions between the trained and untrained corpus. Additionally, we observe that traditional channel normalization techniques are not effective in improving cross-corpus generalization. Further, we evaluate publicly available datasets that are promising for generalization. We find one particular corpus to be significantly better than others. Finally, we find that using a smaller frame shift in short-time processing of speech can significantly improve cross-corpus generalization. The proposed techniques to address cross-corpus generalization include channel normalization, better training corpus, and smaller frame shift in short-time Fourier transform (STFT). These techniques together improve the objective intelligibility and quality scores on untrained corpora significantly. Ashutosh Pandey 0004, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech EnhancementabstractDeep neural network (DNN) embeddings for speaker recognition have recently attracted much attention. Compared to i-vectors, they are more robust to noise and room reverberation as DNNs leverage large-scale training. This article addresses the question of whether speech enhancement approaches are still useful when DNN embeddings are used for speaker recognition. We investigate single- and multi-channel speech enhancement for text-independent speaker verification based on x-vectors in conditions where strong diffuse noise and reverberation are both present. Single-channel (monaural) speech enhancement is based on complex spectral mapping and is applied to individual microphones. We use masking-based minimum variance distortion-less response (MVDR) beamformer and its rank-1 approximation for multi-channel speech enhancement. We propose a novel method of deriving time-frequency masks from the estimated complex spectrogram. In addition, we investigate gammatone frequency cepstral coefficients (GFCCs) as robust speaker features. Systematic evaluations and comparisons on the NIST SRE 2010 retransmitted corpus show that both monaural and multi-channel speech enhancement significantly outperform x-vector's performance, and our covariance matrix estimate is effective for the MVDR beamformer. Hassan Taherian, Zhongqiu Wang 0001, Jorge Chang, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Learning Complex Spectral Mapping With Gated Convolutional Recurrent Networks for Monaural Speech EnhancementabstractPhase is important for perceptual quality of speech. However, it seems intractable to directly estimate phase spectra through supervised learning due to their lack of spectrotemporal structure in it. Complex spectral mapping aims to estimate the real and imaginary spectrograms of clean speech from those of noisy speech, which simultaneously enhances magnitude and phase responses of speech. Inspired by multi-task learning, we propose a gated convolutional recurrent network (GCRN) for complex spectral mapping, which amounts to a causal system for monaural speech enhancement. Our experimental results suggest that the proposed GCRN substantially outperforms an existing convolutional neural network (CNN) for complex spectral mapping in terms of both objective speech intelligibility and quality. Moreover, the proposed approach yields significantly higher STOI and PESQ than magnitude spectral mapping and complex ratio masking. We also find that complex spectral mapping with the proposed GCRN provides an effective phase estimate. Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Bridging the Gap Between Monaural Speech Enhancement and Recognition With Distortion-Independent Acoustic ModelingabstractMonaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago. Although enhanced speech has been demonstrated to have better intelligibility and quality for human listeners, feeding it directly to automatic speech recognition (ASR) systems trained with noisy speech has not produced expected improvements in ASR performance. The lack of an enhancement benefit on recognition, or the gap between monaural speech enhancement and recognition, is often attributed to speech distortions introduced in the enhancement process. In this article, we analyze the distortion problem, compare different acoustic models, and investigate a distortion-independent training scheme for monaural speech recognition. Experimental results suggest that distortion-independent acoustic modeling is able to overcome the distortion problem. Such an acoustic model can also work with speech enhancement models different from the one used during training. Moreover, the models investigated in this paper outperform the previous best system on the CHiME-2 corpus. Ke Tan 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Deep Learning Based Target Cancellation for Speech DereverberationabstractThis study investigates deep learning based single- and multi-channel speech dereverberation. For single-channel processing, we extend magnitude-domain masking and mapping based dereverberation to complex-domain mapping, where deep neural networks (DNNs) are trained to predict the real and imaginary (RI) components of the direct-path signal from reverberant (and noisy) ones. For multi-channel processing, we first compute a minimum variance distortionless response (MVDR) beamformer to cancel the direct-path signal, and then feed the RI components of the cancelled signal, which is expected to be a filtered version of non-target signals, as additional features to perform dereverberation. Trained on a large dataset of simulated room impulse responses, our models show excellent speech dereverberation and recognition performance on the test set of the REVERB challenge, consistently better than single- and multi-channel weighted prediction error (WPE) algorithms. Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Complex Spectral Mapping for Single- and Multi-Channel Speech Enhancement and Robust ASRabstractThis study proposes a complex spectral mapping approach for single- and multi-channel speech enhancement, where deep neural networks (DNNs) are used to predict the real and imaginary (RI) components of the direct-path signal from noisy and reverberant ones. The proposed system contains two DNNs. The first one performs single-channel complex spectral mapping. The estimated complex spectra are used to compute a minimum variance distortion-less response (MVDR) beamformer. The RI components of beamforming results, which encode spatial information, are then combined with the RI components of the mixture to train the second DNN for multi-channel complex spectral mapping. With estimated complex spectra, we also propose a novel method of time-varying beamforming. State-of-the-art performance is obtained on the speech enhancement and recognition tasks of the CHiME-4 corpus. More specifically, our system obtains 6.82%, 3.19% and 2.00% word error rates (WER) respectively on the single-, two-, and six-microphone tasks of CHiME-4, significantly surpassing the current best results of 9.15%, 3.91% and 2.24% WER. Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Monaural Speech Dereverberation Using Temporal Convolutional Networks With Self AttentionabstractIn daily listening environments, human speech is often degraded by room reverberation, especially under highly reverberant conditions. Such degradation poses a challenge for many speech processing systems, where the performance becomes much worse than in anechoic environments. To combat the effect of reverberation, we propose a monaural (single-channel) speech dereverberation algorithm using temporal convolutional networks with self attention. Specifically, the proposed system includes a self-attention module to produce dynamic representations given input features, a temporal convolutional network to learn a nonlinear mapping from such representations to the magnitude spectrum of anechoic speech, and a one-dimensional (1-D) convolution module to smooth the enhanced magnitude among adjacent frames. Systematic evaluations demonstrate that the proposed algorithm improves objective metrics of speech quality in a wide range of reverberant conditions. In addition, it generalizes well to untrained reverberation times, room sizes, measured room impulse responses, real-world recorded noisy-reverberant speech, and different speakers. Yan Zhao 0010, DeLiang Wang, Buye Xu, Tao Zhang 0024 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | TCNN: Temporal Convolutional Neural Network for Real-time Speech Enhancement in the Time DomainabstractThis work proposes a fully convolutional neural network (CNN) for real-time speech enhancement in the time domain. The proposed CNN is an encoder-decoder based architecture with an additional temporal convolutional module (TCM) inserted between the encoder and the decoder. We call this architecture a Temporal Convolutional Neural Network (TCNN). The encoder in the TCNN creates a low dimensional representation of a noisy input frame. The TCM uses causal and dilated convolutional layers to utilize the encoder output of the current and previous frames. The decoder uses the TCM output to reconstruct the enhanced frame. The proposed model is trained in a speaker- and noise-independent way. Experimental results demonstrate that the proposed model gives consistently better enhancement results than a state-of-the-art real-time convolutional recurrent model. Moreover, since the model is fully convolutional, it has much fewer trainable parameters than earlier models. Ashutosh Pandey 0004, DeLiang Wang |
ICASSP | 2 |
| 2019 | Exploring Deep Complex Networks for Complex Spectrogram EnhancementabstractA recent study has demonstrated the effectiveness of complex-valued deep neural networks (CDNNs) using newly developed tools such as complex batch normalization and complex residual blocks. Motivated by the fact that CDNNs are well suited for the processing of complex-domain representations, we explore CDNNs for speech enhancement. In particular, we train a CDNN that learns to map the complex-valued noisy short-time Fourier transform (STFT) to the clean STFT. Additionally, we propose the complex-valued extensions of the parametric rectified linear unit (PReLU) nonlinearity that helps to improve the performance of CDNN. Experimental results demonstrate that a CDNN using the proposed nonlinearity can give similar or better enhancement results compared to real-valued deep neural networks (DNNs). Ashutosh Pandey 0004, DeLiang Wang |
ICASSP | 2 |
| 2019 | Complex Spectral Mapping with a Convolutional Recurrent Network for Monaural Speech EnhancementabstractPhase is important for perceptual quality in speech enhancement. However, it seems intractable to directly estimate phase spectrogram through supervised learning due to lack of clear structure in phase spectrogram. Complex spectral mapping aims to estimate the real and imaginary spectrograms of clean speech from those of noisy speech, which simultaneously enhances magnitude and phase responses of noisy speech. In this paper, we propose a new convolutional recurrent network (CRN) for complex spectral mapping, which leads to a causal system for noise- and speaker-independent speech enhancement. In terms of objective intelligibility and perceptual quality, the proposed CRN significantly outperforms an existing convolutional neural network (CNN) for complex spectral mapping, as well as a strong CRN for magnitude spectral mapping. We additionally incorporate a newly-developed group strategy to substantially reduce the number of trainable parameters and the computational cost without sacrificing performance. Ke Tan 0001, DeLiang Wang |
ICASSP | 2 |
| 2019 | Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk ScenariosabstractIn mobile speech communication, the quality and intelligibility of the received speech can be severely degraded by background noise if the far-end talker is in an adverse acoustic environment. Therefore, speech enhancement algorithms are typically integrated into mobile phones to remove background noise. In this paper, we propose a novel deep learning based framework for real-time speech enhancement on dual-microphone mobile phones in a close-talk scenario. It incorporates a convolutional recurrent network (CRN) with high computational efficiency. In addition, the framework amounts to a causal system, which is necessary for real-time processing on mobile phones. We find that the proposed approach consistently outperforms a deep neural network (DNN) based method, as well as two traditional methods for speech enhancement. Ke Tan 0001, Xueliang Zhang 0001, DeLiang Wang |
ICASSP | 3 |
| 2019 | Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric PerspectiveabstractThis study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture oftwo sources, with their magnitudes accurately estimated and under a geometric constraint, the absolute phase difference between each source and the mixture can be uniquely determined; in addition, the source phases at each time-frequency T - F unit can be narrowed down to only two candidates. To pick the right candidate, we propose three algorithms based on iterative phase reconstruction, group delay estimation, and phase-difference sign prediction. State-of-the-art results are obtained on the publicly available wsj0-2mix and 3 mix corpus. Zhongqiu Wang 0001, Ke Tan 0001, DeLiang Wang |
ICASSP | 3 |
| 2019 | Robust Sparse Multichannel Active Noise ControlabstractMultichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ1-norm constraint to the standard MC-ANC helps to reduce the complexity of the system and accelerate the convergence rate. However, the convergence performance of ℓ1-norm constrained MC-ANC (cℓ1-MC-ANC) degrades significantly in reverberant environments. In this paper, we analyze the necessity of using sparsity-inducing algorithms with distinct zero-attracting strengths over loudspeakers, and then derive three algorithms of this kind in the complex domain. Simulation results show that, compared to cℓ1-MC-ANC, the proposed algorithms exhibit faster convergence or higher noise reduction at steady state in both free field and reverberant environments. Jingli Xie, Danqi Jin, Wen Zhang 0002, Xiao-Lei Zhang 0001, Jie Chen 0022, DeLiang Wang |
ICASSP | 6 |
| 2019 | Deep Learning Based Multi-Channel Speaker Recognition in Noisy and Reverberant Environments
Hassan Taherian, Zhongqiu Wang 0001, DeLiang Wang |
INTERSPEECH | 3 |
| 2019 | Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic ModelingabstractMonaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago.Although enhanced speech has been demonstrated to have better intelligibility and quality for human listeners, feeding it directly to automatic speech recognition (ASR) systems trained with noisy speech has not produced expected improvements in ASR performance.The lack of an enhancement benefit on recognition, or the gap between monaural speech enhancement and recognition, is often attributed to speech distortions introduced in the enhancement process.In this study, we analyze the distortion problem, compare different acoustic models, and investigate a distortionindependent training scheme for monaural speech recognition.Experimental results suggest that distortion-independent acoustic modeling is able to overcome the distortion problem.Such an acoustic model can also work with speech enhancement models different from the one used during training.Moreover, the models investigated in this paper outperform the previous best system on the CHiME-2 corpus. Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 3 |
| 2019 | Enhanced Spectral Features for Distortion-Independent Acoustic Modeling
DeLiang Wang |
INTERSPEECH | 2 |
| 2019 | Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distortions
Hao Zhang 0112, Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 3 |
| 2019 | Deep Learning for Talker-Dependent Reverberant Speaker Separation: An Empirical StudyabstractSpeaker separation refers to the problem of separating speech signals from a mixture of simultaneous speakers. Previous studies are limited to addressing the speaker separation problem in anechoic conditions. This paper addresses the problem of talker-dependent speaker separation in reverberant conditions, which are characteristic of real-world environments. We employ recurrent neural networks with bidirectional long short-term memory (BLSTM) to separate and dereverberate the target speech signal. We propose two-stage networks to effectively deal with both speaker separation and speech dereverberation. In the two-stage model, the first stage separates and dereverberates two-talker mixtures and the second stage further enhances the separated target signal. We have extensively evaluated the two-stage architecture, and our empirical results demonstrate large improvements over unprocessed mixtures and clear performance gain over single-stage networks in a wide range of target-to-interferer ratios and reverberation times in simulated as well as recorded rooms. Moreover, we show that time-frequency masking yields better performance than spectral mapping for reverberant speaker separation. Masood Delfarah, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Divide and Conquer: A Deep CASA Approach to Talker-Independent Monaural Speaker SeparationabstractWe address talker-independent monaural speaker separation from the perspectives of deep learning and computational auditory scene analysis (CASA). Specifically, we decompose the multi-speaker separation task into the stages of simultaneous grouping and sequential grouping. Simultaneous grouping is first performed in each time frame by separating the spectra of different speakers with a permutation-invariantly trained neural network. In the second stage, the frame-level separated spectra are sequentially grouped to different speakers by a clustering network. The proposed deep CASA approach optimizes frame-level separation and speaker tracking in turn, and produces excellent results for both objectives. Experimental results on the benchmark WSJ0-2mix database show that the new approach achieves the state-of-the-art results with a modest model size. DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | A New Framework for CNN-Based Speech Enhancement in the Time DomainabstractThis paper proposes a new learning mechanism for a fully convolutional neural network (CNN) to address speech enhancement in the time domain. The CNN takes as input the time frames of noisy utterance and outputs the time frames of the enhanced utterance. At the training time, we add an extra operation that converts the time domain to the frequency domain. This conversion corresponds to simple matrix multiplication, and is hence differentiable implying that a frequency domain loss can be used for training in the time domain. We use mean absolute error loss between the enhanced short-time Fourier transform (STFT) magnitude and the clean STFT magnitude to train the CNN. This way, the model can exploit the domain knowledge of converting a signal to the frequency domain for analysis. Moreover, this approach avoids the well-known invalid STFT problem since the proposed CNN operates in the time domain. Experimental results demonstrate that the proposed method substantially outperforms the other methods of speech enhancement. The proposed method is easy to implement and applicable to related speech processing tasks that require time-frequency masking or spectral mapping. Ashutosh Pandey 0004, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Gated Residual Networks With Dilated Convolutions for Monaural Speech EnhancementabstractFor supervised speech enhancement, contextual information is important for accurate mask estimation or spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term contexts for tracking a target speaker, we treat speech enhancement as a sequence-to-sequence mapping, and present a novel convolutional neural network (CNN) architecture for monaural speech enhancement. The key idea is to systematically aggregate contexts through dilated convolutions, which significantly expand receptive fields. The CNN model additionally incorporates gating mechanisms and residual learning. Our experimental results suggest that the proposed model generalizes well to untrained noises and untrained speakers. It consistently outperforms a DNN, a unidirectional long short-term memory (LSTM) model and a bidirectional LSTM model in terms of objective speech intelligibility and quality metrics. Moreover, the proposed model has far fewer parameters than DNN and LSTM models. Ke Tan 0001, Jitong Chen, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Combining Spectral and Spatial Features for Deep Learning Based Blind Speaker SeparationabstractThis study tightly integrates complementary spectral and spatial features for deep learning based multi-channel speaker separation in reverberant environments. The key idea is to localize individual speakers so that an enhancement network can be trained on spatial as well as spectral features to extract the speaker from an estimated direction and with specific spectral structures. The spatial and spectral features are designed in a way such that the trained models are blind to the number of microphones and microphone geometry. To determine the direction of the speaker of interest, we identify time-frequency (T-F) units dominated by that speaker and only use them for direction estimation. The T-F unit level speaker dominance is determined by a two-channel chimera++ network, which combines deep clustering and permutation invariant training at the objective function level, and integrates spectral and interchannel phase patterns at the input feature level. In addition, T-F masking based beamforming is tightly integrated in the system by leveraging the magnitudes and phases produced by beamforming. Strong separation performance has been observed on reverberant talker-independent speaker separation, which separates reverberant speaker mixtures based on a random number of microphones arranged in arbitrary linear-array geometry. Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Robust Speaker Localization Guided by Deep Learning-Based Time-Frequency MaskingabstractDeep learning-based time-frequency (T-F) masking has dramatically advanced monaural (single-channel) speech separation and enhancement. This study investigates its potential for direction of arrival (DOA) estimation in noisy and reverberant environments. We explore ways of combining T-F masking and conventional localization algorithms, such as generalized cross correlation with phase transform, as well as newly proposed algorithms based on steered-response SNR and steering vectors. The key idea is to utilize deep neural networks (DNNs) to identify speech dominant T-F units containing relatively clean phase for DOA estimation. Our DNN is trained using only monaural spectral information, and this makes the trained model directly applicable to arrays with various numbers of microphones arranged in diverse geometries. Although only monaural information is used for training, experimental results show strong robustness of the proposed approach in new environments with intense noise and room reverberation, outperforming traditional DOA estimation methods by large margins. Our study also suggests that the ideal ratio mask and its variants remain effective training targets for robust speaker localization. Zhongqiu Wang 0001, Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Two-Stage Deep Learning for Noisy-Reverberant Speech EnhancementabstractIn real-world situations, speech reaching our ears is commonly corrupted by both room reverberation and background noise. These distortions are detrimental to speech intelligibility and quality, and also pose a serious problem to many speech-related applications, including automatic speech and speaker recognition. In order to deal with the combined effects of noise and reverberation, we propose a two-stage strategy to enhance corrupted speech, where denoising and dereverberation are conducted sequentially using deep neural networks. In addition, we design a new objective function that incorporates clean phase during model training to better estimate spectral magnitudes, which would in turn yield better phase estimates when combined with iterative phase reconstruction. The two-stage model is then jointly trained to optimize the proposed objective function. Systematic evaluations and comparisons show that the proposed algorithm improves objective metrics of speech intelligibility and quality substantially, and significantly outperforms previous one-stage enhancement systems. Yan Zhao 0010, Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Recurrent Neural Networks for Cochannel Speech Separation in Reverberant EnvironmentsabstractSpeech separation is a fundamental problem in speech and signal processing. A particular challenge is monaural separation of cochannel speech, or a two-talker mixture, in a reverberant environment. In this paper, we study recurrent neural networks (RNNs) with long short-term memory (LSTM) in separating and enhancing speech signals in reverberant cochannel mixtures. Our investigation shows that RNNs are effective in separating reverberant speech signals. In addition, RNNs significantly outperform deep feedforward networks based on objective speech intelligibility and quality measures. We also find that the best performance is achieved when the ideal ratio mask (IRM) is used as the training target in comparison with alternative training targets. While trained using reverberant signals generated by simulated room impulse responses (RIRs), our model generalizes well to conditions where the signals are generated by recorded RIRs. Masood Delfarah, DeLiang Wang |
ICASSP | 2 |
| 2018 | A Casa Approach to Deep Learning Based Speaker-Independent Co-Channel Speech SeparationabstractWe address speaker-independent co-channel speech separation from the computational auditory scene analysis (CAS A) perspective. Specifically, we decompose the two-speaker separation task into the stages of simultaneous grouping and sequential grouping. Simultaneous grouping is first performed at the frame level by separating the spectra of two speakers with a permutation-invariantly trained recurrent neural network (RNN). In the second stage, the simultaneously separated spectra at each frame are sequentially grouped into the utterances of the two underlying speakers by a clustering RNN. Overall optimization is then performed to fine tune the two-stage system. The proposed CASA approach takes advantage of permutation invariant training (PIT) and deep clustering (DC), but overcomes their shortcomings. Experiments show that the proposed system improves over the best reported results of PIT and DC. DeLiang Wang |
ICASSP | 2 |
| 2018 | Permutation Invariant Training for Speaker-Independent Multi-Pitch TrackingabstractSpeaker-independent multi-pitch tracking has been a long-standing problem in speech processing. In this study, we extend a recurrent neural network-factorial hidden Markov model (RNN-FHMM) framework and use the utterance-level permutation invariant training (uPIT) criterion for multi-pitch tracking. Separated speech and label permutations from a speech separation uPIT-RNN have been further incorporated to improve pitch tracking performance. We evaluate our methods on the GRID database. Results indicate that the proposed speech separation-pitch tracking system with matched uPIT label permutations outperforms all other gender-dependent and speaker-independent multi-pitch trackers. The improvement is more significant for challenging same-gender mixtures. DeLiang Wang |
ICASSP | 2 |
| 2018 | On Adversarial Training and Loss Functions for Speech EnhancementabstractGenerative adversarial networks (GANs) are becoming increasingly popular for image processing tasks. Researchers have started using GAN s for speech enhancement, but the advantage of using the GAN framework has not been established for speech enhancement. For example, a recent study reports encouraging enhancement results, but we find that the architecture of the generator used in the GAN gives better performance when it is trained alone using the L1loss. This work presents a new GAN for speech enhancement, and obtains performance improvement with the help of adversarial training. A deep neural network (DNN) is used for time-frequency mask estimation, and it is trained in two ways: regular training with the L1loss and training using the GAN framework with the help of an adversary discriminator. Experimental results suggest that the GAN framework improves speech enhancement performance. Further exploration of loss functions, for speech enhancement, suggests that the L1loss is consistently better than the L2loss for improving the perceptual quality of noisy speech. Ashutosh Pandey 0004, DeLiang Wang |
ICASSP | 2 |
| 2018 | Gated Residual Networks with Dilated Convolutions for Supervised Speech SeparationabstractIn supervised speech separation, deep neural networks (DNNs) are typically employed to predict an ideal time-frequency (T-F) mask in order to remove background interference. However, the performance of DNNs is frequently degraded for untrained noises and speakers. Inspired by recent research on dilated convolutions for context aggregation, we propose a novel convolutional neural network (CNN) to deal with noise- and speaker-independent speech separation. The proposed model incorporates dilated convolutions, gating mechanisms and residual learning. We find that the proposed model consistently outperforms a state-of-the-art long short-term memory (LSTM) based model in terms of objective speech intelligibility and quality. Additionally, the proposed CNN is more computationally efficient than the LSTM model. Ke Tan 0001, Jitong Chen, DeLiang Wang |
ICASSP | 3 |
| 2018 | Utterance-Wise Recurrent Dropout and Iterative Speaker Adaptation for Robust Monaural Speech RecognitionabstractThis study addresses monaural (single-microphone) automatic speech recognition (ASR) in adverse acoustic conditions. Our study builds on a state-of-the-art monaural robust ASR method that uses a wide residual network with bidirectional long short-term memory (BLSTM). We propose a novel utterance-wise dropout method for training LSTM networks and an iterative speaker adaptation technique. When evaluated on the monaural speech recognition task of the CHiME-4 corpus, our model yields a word error rate (WER) of 8.28% using the baseline language model, outperforming the previous best monaural ASR by 16.19% relatively. DeLiang Wang |
ICASSP | 2 |
| 2018 | Filter-and-Convolve: A Cnn Based Multichannel Complex Concatenation Acoustic ModelabstractWe propose a convolutional neural network (CNN) based multichannel complex-domain concatenation acoustic model. The proposed model extracts speech-specific information from multichannel noisy speech signals. In addition, we design two CNN templates that have wide applicability and several speaker adaptation methods for the multichannel complex concatenation acoustic model. Even with a simple BeamformIt beamformer and the baseline language model, our method obtains a word error rate (WER) of 5.39% on the CHiME-4 corpus, outperforming the previous best result by 13.06% relatively. Using an MVDR beamformer, our model outperforms the corresponding best system by 9.77% relatively. DeLiang Wang |
ICASSP | 2 |
| 2018 | Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASRabstractDeep learning based single-channel time-frequency (T-F) masking has shown considerable potential for beamforming and robust ASR. This paper proposes a simple but novel relative transfer function (RTF) estimation algorithm for microphone arrays, where the RTF between a reference signal and a non-reference signal at each frequency band is estimated as a weighted average of the ratios of the two STFT (short-time Fourier transform) coefficients of the speech-dominant T-F units. Similarly, the noise covariance matrix is estimated from noise-dominant T-F units. An MVDR beamformer is then constructed for robust ASR. Experiments on the two- and six-channel track of the CHiME-4 challenge show consistent improvement over a weighted delay-and-sum (WDAS) beamformer, a generalized eigenvector beamformer, a parameterized multi-channel Wiener filter, an MVDR beamformer based on conventional direction of arrival (DOA) estimation, and two MVDR beamformers both based on eigendecomposition. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2018 | On Spatial Features for Supervised Speech Separation and its Application to Beamforming and Robust ASRabstractThis study integrates complementary spectral and spatial information to elevate deep learning based time-frequency masking and acoustic beamforming. Coherence and directional features are designed as additional input features for deep neural network training to remove diffuse noise and other directional interferences pervasive in real-world recordings. The diffuse and directional features are designed to be relatively invariant to the underlying target direction, number of microphones and microphone geometry. The estimated masks are then utilized to compute steering vectors and spatial covariance matrices for beamforming and robust ASR. Experiments on the CHiME-4 dataset demonstrate the effectiveness of the proposed approach. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2018 | Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term MemoryabstractHuman speech is usually distorted by room reverberation. These corruptions degrade speech quality and intelligibility, especially under a long reverberation time, and they also pose a serious problem for many speech-related applications such as automatic speech recognition. In this paper, we propose a supervised speech dereverberation algorithm that models late reverberation using a recurrent neural network (RNN) with long short-term memory (LSTM). By taking advantage of LSTM's ability to capture a long history, late reverberation can be effectively removed by the proposed approach. Systematic evaluations indicate that our approach improves the quality of reverberant speech in a wide range of reverberant conditions. Moreover, the proposed system is a causal system, which can be applied in real-time applications. Yan Zhao 0010, DeLiang Wang, Buye Xu, Tao Zhang 0024 |
ICASSP | 2 |
| 2018 | A New Framework for Supervised Speech Enhancement in the Time Domain
Ashutosh Pandey 0004, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | A Convolutional Recurrent Neural Network for Real-Time Speech Enhancement
Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | A Two-Stage Approach to Noisy Cochannel Speech Separation with Gated Residual Networks
Ke Tan 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | End-to-End Speech Separation with Unfolded Iterative Phase ReconstructionabstractThis paper proposes an end-to-end approach for single-channel speaker-independent multi-speaker speech separation, where time-frequency (T-F) masking, the short-time Fourier transform (STFT), and its inverse are represented as layers within a deep network.Previous approaches, rather than computing a loss on the reconstructed signal, used a surrogate loss based on the target STFT magnitudes.This ignores reconstruction error introduced by phase inconsistency.In our approach, the loss function is directly defined on the reconstructed signals, which are optimized for best separation.In addition, we train through unfolded iterations of a phase reconstruction algorithm, represented as a series of STFT and inverse STFT layers.While mask values are typically limited to lie between zero and one for approaches using the mixture phase for reconstruction, this limitation is less relevant if the estimated magnitudes are to be used together with phase reconstruction.We thus propose several novel activation functions for the output layer of the T-F masking, to allow mask values beyond one.On the publiclyavailable wsj0-2mix dataset, our approach achieves state-ofthe-art 12.6 dB scale-invariant signal-to-distortion ratio (SI-SDR) and 13.1 dB SDR, revealing new possibilities for deep learning based phase reconstruction and representing a fundamental progress towards solving the notoriously-hard cocktail party problem. Zhongqiu Wang 0001, Jonathan Le Roux, DeLiang Wang, John R. Hershey |
INTERSPEECH | 3 |
| 2018 | Integrating Spectral and Spatial Features for Multi-Channel Speaker Separation
Zhongqiu Wang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | All-Neural Multi-Channel Speech Enhancement
Zhongqiu Wang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | Robust TDOA Estimation Based on Time-Frequency Masking and Deep Neural Networks
Zhongqiu Wang 0001, Xueliang Zhang 0001, DeLiang Wang |
INTERSPEECH | 3 |
| 2018 | Deep Learning for Acoustic Echo Cancellation in Noisy and Double-Talk Scenarios
Hao Zhang 0112, DeLiang Wang |
INTERSPEECH | 2 |
| 2018 | Fostering deep learning and beyond
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2018 | Supervised Speech Separation Based on Deep Learning: An OverviewabstractSpeech separation is the task of separating target speech from background interference. Traditionally, speech separation is studied as a signal processing problem. A more recent approach formulates speech separation as a supervised learning problem, where the discriminative patterns of speech, speakers, and background noise are learned from training data. Over the past decade, many supervised separation algorithms have been put forward. In particular, the recent introduction of deep learning to supervised speech separation has dramatically accelerated progress and boosted separation performance. This paper provides a comprehensive overview of the research on deep learning based supervised speech separation in the last several years. We first introduce the background of speech separation and the formulation of supervised separation. Then, we discuss three main components of supervised separation: learning machines, training targets, and acoustic features. Much of the overview is on separation algorithms where we review monaural methods, including speech enhancement (speech-nonspeech separation), speaker separation (multitalker separation), and speech dereverberation, as well as multimicrophone techniques. The important issue of generalization, unique to supervised learning, is discussed. This overview provides a historical perspective on how advances are made. In addition, we discuss a number of conceptual issues, including what constitutes the target source. DeLiang Wang, Jitong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Robust speaker recognition based on DNN/i-vectors and speech separationabstractRecent research shows that the i-vector framework for speaker recognition can significantly benefit from phonetic information. A common approach is to use a deep neural network (DNN) trained for automatic speech recognition to generate a universal background model (UBM). Studies in this area have been done in relatively clean conditions. However, strong background noise is known to severely reduce speaker recognition performance. This study investigates a phonetically-aware i-vector system in noisy conditions. We propose a front-end to tackle the noise problem by performing speech separation and examine its performance for both verification and identification tasks. The proposed separation system trains a DNN to estimate the ideal ratio mask of the noisy speech. The separated speech is then used to extract enhanced features for the i-vector framework. We compare the proposed system against a multi-condition trained baseline and a traditional GMM-UBM i-vector system. Our proposed system provides an absolute average improvement of 8% in identification accuracy and 1.2% in equal error rate. Jorge Chang, DeLiang Wang |
ICASSP | 2 |
| 2017 | Time and frequency domain long short-term memory for noise robust pitch trackingabstractPitch tracking in noisy speech is a challenging task as temporal and spectral patterns of the speech signal are both corrupted. This paper proposes long short-term memory (LSTM) based methods for pitch probability estimation. Two architectures are investigated. The first one is conventional LSTM that utilizes recurrent connections to model pitch dynamics. The second one is two-level time-frequency LSTM, with the first level scanning frequency bands and the second level connecting the first level through time. The Viterbi algorithm then takes the probabilistic output from LSTM to generate continuous pitch contours. Experiments show that both proposed models outperform a deep neural network (DNN) based model in most conditions. Time-frequency LSTM achieves the best performance at negative SNRs. DeLiang Wang |
ICASSP | 2 |
| 2017 | Recurrent deep stacking networks for supervised speech separationabstractSupervised speech separation algorithms seldom utilize output patterns. This study proposes a novel recurrent deep stacking approach for time-frequency masking based speech separation, where the output context is explicitly employed to improve the accuracy of mask estimation. The key idea is to incorporate the estimated masks of several previous frames as additional inputs to better estimate the mask of the current frame. Rather than formulating it as a recurrent neural network (RNN), which is potentially much harder to train, we propose to train a deep neural network (DNN) with implicit deep stacking. The estimated masks of the previous frames are updated only at the end of each DNN training epoch, and then the updated estimated masks provide additional inputs to train the DNN in the next epoch. At the test stage, the DNN makes predictions sequentially in a recurrent fashion. In addition, we propose to use the L1loss for training. Experiments on the CHiME-2 (task-2) dataset demonstrate the effectiveness of our proposed approach. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2017 | Unsupervised speaker adaptation of batch normalized acoustic models for robust ASRabstractBatch normalization is a standard technique for training deep neural networks. In batch normalization, the input of each hidden layer is first mean-variance normalized and then linearly transformed before applying non-linear activation functions. We propose a novel unsupervised speaker adaptation technique for batch normalized acoustic models. The key idea is to adjust the linear transformations previously learned by batch normalization for all the hidden layers according to the first-pass decoding results of the speaker-independent model. With the adjusted linear transformations for each test speaker, the test distribution of the input of each hidden layer better matches the training distribution. Experiments on the CHiME-3 dataset demonstrate the effectiveness of the proposed layer-wise adaptation approach. Our overall system obtains 4.24% WER on the real subset of the test data, which represents the best reported result on this dataset to date and a relative 27.3% error reduction over the previous best result. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2017 | Speech dereverberation and denoising using complex ratio masksabstractTraditional speech separation systems enhance the magnitude response of noisy speech. Recent studies, however, have shown that perceptual speech quality is significantly improved when magnitude and phase are both enhanced. These studies, however, have not determined if phase enhancement is beneficial in environments that contain reverberation as well as noise. In this paper, we present an approach that jointly enhances the magnitude and phase of reverberant and noisy speech. We use a deep neural network to estimate the real and imaginary components of the complex ideal ratio mask (cIRM), which results in clean and anechoic speech when applied to a reverberant-noisy mixture. Our results show that phase is important for dereverberation, and that complex ratio masking outperforms related methods. Donald S. Williamson, DeLiang Wang |
ICASSP | 2 |
| 2017 | A speech enhancement algorithm by iterating single- and multi-microphone processing and its application to robust ASRabstractWe propose a speech enhancement algorithm based on single- and multi-microphone processing techniques. The core of the algorithm estimates a time-frequency mask which represents the target speech and use masking-based beamforming to enhance corrupted speech. Specifically, in single-microphone processing, the received signals of a microphone array are treated as individual signals and we estimate a mask for the signal of each microphone using a deep neural network (DNN). With these masks, in multi-microphone processing, we calculate a spatial covariance matrix of noise and steering vector for beamforming. In addition, we propose a masking-based post-filter to further suppress the noise in the output of beamforming. Then, the enhanced speech is sent back to DNN for mask re-estimation. When these steps are iterated for a few times, we obtain the final enhanced speech. The proposed algorithm is evaluated as a frontend for automatic speech recognition (ASR) and achieves a 5.05% average word error rate (WER) on the real environment test set of CHiME-3, outperforming the current best algorithm by 13.34%. Xueliang Zhang 0001, Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 3 |
| 2017 | A two-stage algorithm for noisy and reverberant speech enhancementabstractIn daily listening environments, speech is commonly corrupted by room reverberation and background noise. These distortions are detrimental to speech intelligibility and quality, and also severely degrade the performance of automatic speech and speaker recognition systems. In this paper, we propose a two-stage algorithm to deal with the confounding effects of noise and reverberation separately, where denoising and dereverberation are conducted sequentially using deep neural networks. In addition, we design a new objective function that incorporates clean phase information during training. As the objective function emphasizes more important time-frequency (T-F) units, better estimated magnitude is obtained during testing. By jointly training the two-stage model to optimize the proposed objective function, our algorithm improves objective metrics of speech intelligibility and quality significantly, and substantially outperforms one-stage enhancement baselines. Yan Zhao 0010, Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 3 |
| 2017 | Binaural Reverberant Speech Separation Based on Deep Neural Networks
Xueliang Zhang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2017 | Promoting Further Developments of Neural Networks
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2017 | Announcement of the Neural Networks Best Paper Award
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2017 | Features for Masking-Based Monaural Speech Separation in Reverberant ConditionsabstractMonaural speech separation is a fundamental problem in speech and signal processing. This problem can be approached from a supervised learning perspective by predicting an ideal time-frequency mask from features of noisy speech. In reverberant conditions at low signal-to-noise ratios (SNRs), accurate mask prediction is challenging and can benefit from effective features. In this paper, we investigate an extensive set of acoustic-phonetic features extracted in adverse conditions. Deep neural networks are used as the learning machine, and separation performance is evaluated using standard objective speech intelligibility metrics. Separation performance is systematically evaluated in both nonspeech and speech interference, in a variety of SNRs, reverberation times, and direct-to-reverberant energy ratios. Considerable performance improvement is observed by using contextual information, likely due to temporal effects of room reverberation. In addition, we construct feature combination sets using a sequential floating forward selection algorithm, and combined features outperform individual ones. We also find that optimal feature sets in anechoic conditions are different from those in reverberant conditions. Masood Delfarah, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Time-Frequency Masking in the Complex Domain for Speech Dereverberation and DenoisingabstractIn real-world situations, speech is masked by both background noise and reverberation, which negatively affect perceptual quality and intelligibility. In this paper, we address monaural speech separation in reverberant and noisy environments. We perform dereverberation and denoising using supervised learning with a deep neural network. Specifically, we enhance the magnitude and phase by performing separation with an estimate of the complex ideal ratio mask. We define the complex ideal ratio mask so that direct speech results after the mask is applied to reverberant and noisy speech. Our approach is evaluated using simulated and real room impulse responses, and with background noises. The proposed approach improves objective speech quality and intelligibility significantly. Evaluations and comparisons show that it outperforms related methods in many reverberant and noisy environments. Donald S. Williamson, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Deep Learning Based Binaural Speech Separation in Reverberant EnvironmentsabstractSpeech signal is usually degraded by room reverberation and additive noises in real environments. This paper focuses on separating target speech signal in reverberant conditions from binaural inputs. Binaural separation is formulated as a supervised learning problem, and we employ deep learning to map from both spatial and spectral features to a training target. With binaural inputs, we first apply a fixed beamformer and then extract several spectral features. A new spatial feature is proposed and extracted to complement the spectral features. The training target is the recently suggested ideal ratio mask. Systematic evaluations and comparisons show that the proposed system achieves very good separation performance and substantially outperforms related algorithms under challenging multi-source and reverberant environments. Xueliang Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Robust pitch tracking in noisy speech using speaker-dependent deep neural networksabstractA reliable estimate of pitch in noisy speech is crucial for many speech applications. In this paper, we propose to use speaker-dependent (SD) deep neural networks (DNNs) to model the harmonic patterns of each speaker. Specifically, SD-DNNs take spectral features as input and estimate probabilistic pitch states at each time frame. We investigate two methods for SD-DNN training. The first one is direct training when speaker-dependent data is sufficient. The second one is speaker adaptation of a speaker-independent (SI) DNN with limited data. The Viterbi algorithm is then used to track pitch through time. Experiments show that both training methods of SD-DNNs outperform an SI-DNN based system as well as a state-of-the-art pitch tracking algorithm in all SNR conditions. DeLiang Wang |
ICASSP | 2 |
| 2016 | Robust speech recognition from ratio masksabstractRobustness against noise is crucial for automatic speech recognition systems in real-world environments. In this paper, we propose a novel approach that performs robust ASR by directly recognizing ratio masks. In the proposed approach, a deep neural network (DNN) is first trained to estimate the ideal ratio mask (IRM) from a noisy utterance and then a convolutional neural network (CNN) is employed to recognize estimated IRMs. The proposed approach has been evaluated on the TIDigits corpus, and the results demonstrate that direct recognition of ratio masks outperforms direct recognition of binary masks and traditional MMSE-HMM based method for robust ASR. Zhongqiu Wang 0001, DeLiang Wang |
ICASSP | 2 |
| 2016 | Phoneme-specific speech separationabstractSpeech separation or enhancement algorithms seldom exploit information about phoneme identities. In this study, we propose a novel phoneme-specific speech separation method. Rather than training a single global model to enhance all the frames, we train a separate model for each phoneme to process its corresponding frames. A robust ASR system is employed to identify the phoneme identity of each frame. This way, the information from ASR systems and language models can directly influence speech separation by selecting a phoneme-specific model to use at the test stage. In addition, phoneme-specific models have fewer variations to model and do not exhibit the data imbalance problem. The improved enhancement results can in turn help recognition. Experiments on the corpus of the second CHiME speech separation and recognition challenge (task-2) demonstrate the effectiveness of this method in terms of objective measures of speech intelligibility and quality, as well as recognition performance. Zhongqiu Wang 0001, Yan Zhao 0010, DeLiang Wang |
ICASSP | 3 |
| 2016 | Complex ratio masking for joint enhancement of magnitude and phaseabstractThe phase response of noisy speech has largely been ignored, but recent research shows the importance of phase for perceptual speech quality. A few phase enhancement approaches have been developed. These systems, however, require a separate algorithm for enhancing the magnitude response. In this paper, we present a novel framework for performing monaural speech separation in the complex domain. We show that much structure is exhibited in the real and imaginary components of the short-time Fourier transform, making the complex domain appropriate for supervised estimation. Consequently, we define the complex ideal ratio mask (cIRM) that jointly enhances the magnitude and phase of noisy speech. We then employ a single deep neural network to estimate both the real and imaginary components of the cIRM. The evaluation results show that complex ratio masking yields high quality speech enhancement, and outperforms related methods that operate in the magnitude domain or separately enhance magnitude and phase. Donald S. Williamson, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2016 | DNN-based enhancement of noisy and reverberant speechabstractIn the real world, speech is usually distorted by both reverberation and background noise. In such conditions, speech intelligibility is degraded substantially, especially for hearing-impaired (HI) listeners. As a consequence, it is essential to enhance speech in the noisy and reverberant environment. Recently, deep neural networks have been introduced to learn a spectral mapping to enhance corrupted speech, and shown significant improvements in objective metrics and automatic speech recognition score. However, listening tests have not yet shown any speech intelligibility benefit. In this paper, we propose to enhance the noisy and reverberant speech by learning a mapping to reverberant target speech rather than anechoic target speech. A preliminary listening test was conducted, and the results show that the proposed algorithm is able to improve speech intelligibility of HI listeners in some conditions. Moreover, we develop a masking-based method for denoising and compare it with the spectral mapping method. Evaluation results show that the masking-based method outperforms the mapping-based method. Yan Zhao 0010, DeLiang Wang, Ivo Merks, Tao Zhang 0024 |
ICASSP | 2 |
| 2016 | Long Short-Term Memory for Speaker Generalization in Supervised Speech SeparationabstractSpeech separation can be formulated as learning to estimate a time-frequency mask from acoustic features extracted from noisy speech. For supervised speech separation, generalization to unseen noises and unseen speakers is a critical issue. Although deep neural networks (DNNs) have been successful in noise-independent speech separation, DNNs are limited in modeling a large number of speakers. To improve speaker generalization, a separation model based on long short-term memory (LSTM) is proposed, which naturally accounts for temporal dynamics of speech. Systematic evaluation shows that the proposed model substantially outperforms a DNN-based model on unseen speakers and unseen noises in terms of objective speech intelligibility. Analyzing LSTM internal representations reveals that LSTM captures long-term speech contexts. It is also found that the LSTM model is more advantageous for low-latency speech separation and it, without future frames, performs better than the DNN model with future frames. The proposed model represents an effective approach for speaker- and noise-independent speech separation. Jitong Chen, DeLiang Wang |
INTERSPEECH | 2 |
| 2016 | A Feature Study for Masking-Based Reverberant Speech Separation
Masood Delfarah, DeLiang Wang |
INTERSPEECH | 2 |
| 2016 | State of Neural Networks Is Strong
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2016 | Noise perturbation for supervised speech separation
Jitong Chen, Yuxuan Wang 0002, DeLiang Wang |
Speech Commun. | 3 |
| 2016 | A Joint Training Framework for Robust Automatic Speech RecognitionabstractRobustness against noise and reverberation is critical for ASR systems deployed in real-world environments. In robust ASR, corrupted speech is normally enhanced using speech separation or enhancement algorithms before recognition. This paper presents a novel joint training framework for speech separation and recognition. The key idea is to concatenate a deep neural network (DNN) based speech separation frontend and a DNN-based acoustic model to build a larger neural network, and jointly adjust the weights in each module. This way, the separation frontend is able to provide enhanced speech desired by the acoustic model and the acoustic model can guide the separation frontend to produce more discriminative enhancement. In addition, we apply sequence training to the jointly trained DNN so that the linguistic information contained in the acoustic and language models can be back-propagated to influence the separation frontend at the training stage. To further improve the robustness, we add more noise- and reverberation-robust features for acoustic modeling. At the test stage, utterance-level unsupervised adaptation is performed to adapt the jointly trained network by learning a linear transformation of the input of the separation frontend. The resulting sequence-discriminative jointly-trained multistream system with run-time adaptation achieves 10.63% average word error rate (WER) on the test set of the reverberant and noisy CHiME-2 dataset (task-2), which represents the best performance on this dataset and a 22.75% error reduction over the best existing method. Zhongqiu Wang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Complex Ratio Masking for Monaural Speech SeparationabstractSpeech separation systems usually operate on the short-time Fourier transform (STFT) of noisy speech, and enhance only the magnitude spectrum while leaving the phase spectrum unchanged. This is done because there was a belief that the phase spectrum is unimportant for speech enhancement. Recent studies, however, suggest that phase is important for perceptual quality, leading some researchers to consider magnitude and phase spectrum enhancements. We present a supervised monaural speech separation approach that simultaneously enhances the magnitude and phase spectra by operating in the complex domain. Our approach uses a deep neural network to estimate the real and imaginary components of the ideal ratio mask defined in the complex domain. We report separation results for the proposed method and compare them to related systems. The proposed approach improves over other methods when evaluated with several objective metrics, including the perceptual evaluation of speech quality (PESQ), and a listening test where subjects prefer the proposed approach with at least a 69% rate. Donald S. Williamson, Yuxuan Wang 0002, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Boosting Contextual Information for Deep Neural Network Based Voice Activity DetectionabstractVoice activity detection (VAD) is an important topic in audio signal processing. Contextual information is important for improving the performance of VAD at low signal-to-noise ratios. Here we explore contextual information by machine learning methods at three levels. At the top level, we employ an ensemble learning framework, named multi-resolution stacking (MRS), which is a stack of ensemble classifiers. Each classifier in a building block inputs the concatenation of the predictions of its lower building blocks and the expansion of the raw acoustic feature by a given window (called a resolution). At the middle level, we describe a base classifier in MRS, named boosted deep neural network (bDNN). bDNN first generates multiple base predictions from different contexts of a single frame by only one DNN and then aggregates the base predictions for a better prediction of the frame, and it is different from computationally-expensive boosting methods that train ensembles of classifiers for multiple base predictions. At the bottom level, we employ the multi-resolution cochleagram feature, which incorporates the contextual information by concatenating the cochleagram features at multiple spectrotemporal resolutions. Experimental results show that the MRS-based VAD outperforms other VADs by a considerable margin. Moreover, when trained on a large amount of noise types and a wide range of signal-to-noise ratios, the MRS-based VAD demonstrates surprisingly good generalization performance on unseen test scenarios, approaching the performance with noise-dependent training. Xiao-Lei Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | A Deep Ensemble Learning Method for Monaural Speech SeparationabstractMonaural speech separation is a fundamental problem in robust speech processing. Recently, deep neural network (DNN)-based speech separation methods, which predict either clean speech or an ideal time-frequency mask, have demonstrated remarkable performance improvement. However, a single DNN with a given window length does not leverage contextual information sufficiently, and the differences between the two optimization objectives are not well understood. In this paper, we propose a deep ensemble method, named multicontext networks, to address monaural speech separation. The first multicontext network averages the outputs of multiple DNNs whose inputs employ different window lengths. The second multicontext network is a stack of multiple DNNs. Each DNN in a module of the stack takes the concatenation of original acoustic features and expansion of the soft output of the lower module as its input, and predicts the ratio mask of the target speaker; the DNNs in the same module employ different contexts. We have conducted extensive experiments with three speech corpora. The results demonstrate the effectiveness of the proposed method. We have also compared the two optimization objectives systematically and found that predicting the ideal time-frequency mask is more efficient in utilizing clean training speech, while predicting clean speech is less sensitive to SNR variations. Xiao-Lei Zhang 0001, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A deep neural network for time-domain signal reconstructionabstractSupervised speech separation has achieved considerable success recently. Typically, a deep neural network (DNN) is used to estimate an ideal time-frequency mask, and clean speech is produced by feeding the mask-weighted output to a resynthesizer in a subsequent step. So far, the success of DNN-based separation lies mainly in improving human speech intelligibility. In this work, we propose a new deep network that directly reconstructs the time-domain clean signal through an inverse fast Fourier transform layer. The joint training of speech resynthesis and mask estimation yields improved objective quality while maintaining the objective intelligibility performance. The proposed system significantly outperforms a recent non-negative matrix factorization based separation system in both objective speech intelligibility and quality. Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 2 |
| 2015 | Deep neural networks for estimating speech model activationsabstractThis paper presents an approach for improving the perceptual quality of speech separated from background noise at low signal-to-noise ratios. Our approach uses two stages of deep neural networks, where the first stage estimates the ideal ratio mask that separates speech from noise, and the second stage maps the ratio-masked speech to the clean speech activation matrices that are used for nonnegative matrix factorization (NMF). Supervised NMF systems make assumptions about the relationship between the activation and basic matrices that do not always hold. Other two-stage approaches combining masking with NMF reconstruction do not account for mask estimation errors. We show that the proposed algorithm achieves higher objective speech quality and intelligibility compared to these related methods. Donald S. Williamson, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2015 | Deep neural networks for cochannel speaker identificationabstractSpeaker identification (SID) in cochannel speech, where two speakers are talking simultaneously over a single recording channel, is a challenging problem. Previous studies address this problem in the anechoic environment under the Gaussian mixture model (GMM) framework. On the other hand, cochannel SID in reverberant conditions has not been addressed. This paper studies cochannel SID in both anechoic and reverberant conditions. We explore deep neural networks (DNNs) for cochannel SID and propose a DNN-based recognition system. Evaluation results demonstrate the proposed DNN-based system outperforms the two state-of-the-art cochannel SID systems in both anechoic and reverberant conditions and various target-to-interferer ratios. Xiaojia Zhao, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2015 | Deep neural network based spectral feature mapping for robust speech recognitionabstractAutomatic speech recognition (ASR) systems suffer from performance degradation under noisy and reverberant conditions. In this work, we explore a deep neural network (DNN) based approach for spectral feature mapping from corrupted speech to clean speech. The DNN based mapping substantially reduces interference and produces estimated clean spectral features for ASR training and decoding. We experiment with several different feature mapping approaches and demonstrate that a DNN trained to predict clean log filterbank coefficients from noisy spectrogram directly can be extremely effective. The experiments show that the ASR systems with these cleaned features perform well under joint noisy and reverberant conditions, and achieve the state-of-the-art results on the CHiME-2 corpus with stereo (corrupted and clean) data. Yanzhang He, Deblin Bagchi, Eric Fosler-Lussier, DeLiang Wang |
INTERSPEECH | 5 |
| 2015 | Speaker-dependent multipitch tracking using deep neural networksabstractMultipitch tracking is important for speech and signal processing. However, it is challenging to design an algorithm that achieves accurate pitch estimation and correct speaker assignment at the same time. In this paper, deep neural networks (DNNs) are used to model the probabilistic pitch states of two simultaneous speakers. To capture speaker-dependent information, two types of DNN with different training strategies are proposed. The first is trained for each speaker enrolled in the system (speaker-dependent DNN), and the second is trained for each speaker pair (speaker-pair-dependent DNN). Several extensions, including gender-pair-dependent DNNs, speaker adaptation of gender-pair-dependent DNNs and training with multiple energy ratios, are introduced later to relax constraints. A factorial hidden Markov model (FHMM) then integrates pitch probabilities and generates the most likely pitch tracks with a junction tree algorithm. Experiments show that the proposed methods substantially outperform other speaker-independent and speaker-dependent multipitch trackers on two-speaker mixtures. With multi-ratio training, the proposed methods achieve consistent performance at various energies ratios of the two speakers in a mixture. DeLiang Wang |
INTERSPEECH | 2 |
| 2015 | Joint training of speech separation, filterbank and acoustic model for robust automatic speech recognitionabstractRobustness is crucial for automatic speech recognition systems in real-world environments. Speech enhancement/separation algorithms are normally used to enhance noisy speech before recognition. However, such algorithms typically introduce distortions unseen by acoustic models. In this study, we propose a novel joint training approach to reduce this distortion problem. At the training stage, we first concatenate a speech separation DNN, a filterbank and an acoustic model DNN to form a deeper network, and then jointly train all of them. This way, the separation frontend and filterbank can provide enhanced speech desired by the acoustic model. In addition, the linguistic information contained in the acoustic model can have a positive effect on the frontend and filberbank. Besides the commonly used log mel-spectrogram feature, we also add more robust features for acoustic modeling. Our system obtains 14.1% average word error rate on the noisy and reverberant CHIME-2 corpus (track 2), which outperforms the previous best result by 8.4% relatively. Zhongqiu Wang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2015 | Multi-resolution stacking for speech separation based on boosted DNNabstractRecent progress in speech separation shows that deep neural networks (DNN) based supervised methods can improve the performance in difficult noise conditions and exhibit good generalization to unseen noise scenarios. However, existing approaches do not explore contextual information sufficiently. In this paper, we focus on exploring contextual information using DNN. The proposed method has two parts—a multi-resolution stacking (MRS) framework and a boosted DNN (bDNN) classifier. The MRS framework trains a stack of classifier ensembles, where each classifier in an ensemble concatenates the raw acoustic feature and the outputs of its bottom ensemble as a new feature, and different classifiers in an ensemble work with different window lengths. The bDNN classifier first generates multiple base predictions for a frame from a given window that is centered on the frame and contains multiple neighboring frames, and then aggregates the base predictions for the final prediction. Our experimental comparison with DNN based speech separation in difficult noise scenarios demonstrates the effectiveness of the proposed method in terms of both prediction accuracy and objective speech intelligibility. Xiao-Lei Zhang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2015 | Learning Spectral Mapping for Speech Dereverberation and DenoisingabstractIn real-world environments, human speech is usually distorted by both reverberation and background noise, which have negative effects on speech intelligibility and speech quality. They also cause performance degradation in many speech technology applications, such as automatic speech recognition. Therefore, the dereverberation and denoising problems must be dealt with in daily listening environments. In this paper, we propose to perform speech dereverberation using supervised learning, and the supervised approach is then extended to address both dereverberation and denoising. Deep neural networks are trained to directly learn a spectral mapping from the magnitude spectrogram of corrupted speech to that of clean speech. The proposed approach substantially attenuates the distortion caused by reverberation, as well as background noise, and is conceptually simple. Systematic experiments show that the proposed approach leads to significant improvements of predicted speech intelligibility and quality, as well as automatic speech recognition in reverberant noisy conditions. Comparisons show that our approach substantially outperforms related methods. Yuxuan Wang 0002, DeLiang Wang, William S. Woods, Ivo Merks, Tao Zhang 0024 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Improving Robustness of Deep Neural Network Acoustic Models via Speech Separation and Joint Adaptive TrainingabstractAlthough deep neural network (DNN) acoustic models are known to be inherently noise robust, especially with matched training and testing data, the use of speech separation as a frontend and for deriving alternative feature representations has been shown to improve performance in challenging environments. We first present a supervised speech separation system that significantly improves automatic speech recognition (ASR) performance in realistic noise conditions. The system performs separation via ratio time-frequency masking; the ideal ratio mask (IRM) is estimated using DNNs. We then propose a framework that unifies separation and acoustic modeling via joint adaptive training. Since the modules for acoustic modeling and speech separation are implemented using DNNs, unification is done by introducing additional hidden layers with fixed weights and appropriate network architecture. On the CHiME-2 medium-large vocabulary ASR task, and with log mel spectral features as input to the acoustic model, an independently trained ratio masking frontend improves word error rates by 10.9% (relative) compared to the noisy baseline. In comparison, the jointly trained system improves performance by 14.4%. We also experiment with alternative feature representations to augment the standard log mel features, like the noise and speech estimates obtained from the separation module, and the standard feature set used for IRM estimation. Our best system obtains a word error rate of 15.4% (absolute), an improvement of 4.6 percentage points over the next best result on this corpus. Arun Narayanan, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Cochannel Speaker Identification in Anechoic and Reverberant ConditionsabstractSpeaker identification (SID) in cochannel speech, where two speakers are talking simultaneously over a single recording channel, is a challenging problem. Previous studies address this problem in the anechoic environment under the Gaussian mixture model (GMM) framework. On the other hand, cochannel SID in reverberant conditions has not been addressed. This paper studies cochannel SID in both anechoic and reverberant conditions. We first investigate GMM-based approaches and propose a combined system that integrates two cochannel SID methods. Second, we explore deep neural networks (DNNs) for cochannel SID and propose a DNN-based recognition system. Evaluation results demonstrate that our proposed systems significantly improve SID performance over recent approaches in both anechoic and reverberant conditions and various target-to-interferer ratios. Xiaojia Zhao, Yuxuan Wang 0002, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Factorization-Based Texture SegmentationabstractThis paper introduces a factorization-based approach that efficiently segments textured images. We use local spectral histograms as features, and construct an M × N feature matrix using M-dimensional feature vectors in an N-pixel image. Based on the observation that each feature can be approximated by a linear combination of several representative features, we factor the feature matrix into two matrices--one consisting of the representative features and the other containing the weights of representative features at each pixel used for linear combination. The factorization method is based on singular value decomposition and nonnegative matrix factorization. The method uses local spectral histograms to discriminate region appearances in a computationally efficient way and at the same time accurately localizes region boundaries. The experiments conducted on public segmentation data sets show the promise of this simple yet powerful approach. Jiangye Yuan, DeLiang Wang, Anil M. Cheriyadat |
IEEE Trans. Image Process. | 2 |
| 2014 | A feature study for classification-based speech separation at very low signal-to-noise ratioabstractSpeech separation is a challenging problem at low signal-to-noise ratios (SNRs). Separation can be formulated as a classification problem. In this study, we focus on the SNR level of -5 dB in which speech is generally dominated by background noise. In such a low SNR condition, extracting robust features from a noisy mixture is crucial for successful classification. Using a common neural network classifier, we systematically compare separation performance of many monaural features. In addition, we propose a new feature called Multi-Resolution Cochleagram (MRCG), which is extracted from four cochlea-grams of different resolutions to capture both local information and spectrotemporal context. Comparisons using two non-stationary noises show a range of feature robustness for speech separation with the proposed MRCG performing the best. We also find that ARMA filtering, a post-processing technique previously used for robust speech recognition, improves speech separation performance by smoothing the temporal trajectories of feature dimensions. Jitong Chen, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2014 | Neural networks for supervised pitch tracking in noiseabstractDetermination of pitch in noise is challenging because of corrupted harmonic structure. In this paper, we extract pitch using supervised learning, where probabilistic pitch states are directly learned from noisy speech. We investigate two alternative neural networks modeling the pitch states given observations. The first one is the feedforward deep neural network (DNN), which is trained on static frame-level features. The second one is the recurrent deep neural network (RNN) capable of learning the temporal dynamics trained on sequential frame-level features. Both DNNs and RNNs produce accurate probabilistic outputs of pitch states, which are then connected into pitch contours by Viterbi decoding. Our systematic evaluation shows that the proposed pitch tracking approaches are robust to different noise conditions and significantly outperform current state-of-the-art pitch tracking techniques. DeLiang Wang |
ICASSP | 2 |
| 2014 | Learning spectral mapping for speech dereverberationabstractReverberation distorts human speech and usually has negative effects on speech intelligibility, especially for hearing-impaired listeners. It also causes performance degradation in automatic speech recognition and speaker identification systems. Therefore, the dereverberation problem must be dealt with in daily listening environments. We propose to use deep neural networks (DNNs) to learn a spectral mapping from the reverberant speech to the anechoic speech. The trained DNN produces the estimated spectral representation of the corresponding anechoic speech. We demonstrate that distortion caused by reverberation is substantially attenuated by the DNN whose outputs can be resynthesized to the dereverebrated speech signal. The proposed approach is simple, and our systematic evaluation shows promising dereverberation results, which are significantly better than those of related systems. Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2014 | Joint noise adaptive training for robust automatic speech recognitionabstractWe explore time-frequency masking to improve noise robust automatic speech recognition. Apart from its use as a frontend, we use it for providing smooth estimates of speech and noise which are then passed as additional features to a deep neural network (DNN) based acoustic model. Such a system improves performance on the Aurora-4 dataset by 10.5% (relative) compared to the previous best published results. By formulating separation as a supervised mask estimation problem, we develop a unified DNN framework that jointly improves separation and acoustic modeling. Our final system outperforms the previous best system on CHiME-2 corpus by 22.1% (relative). Arun Narayanan, DeLiang Wang |
ICASSP | 2 |
| 2014 | A structure-preserving training target for supervised speech separationabstractSupervised learning based speech separation has shown considerable success recently. In its simplest form, a discriminative model is trained as a time-frequency masking function, where the training target is an ideal mask. Ideal masks, such as the ideal binary masks, are structured spectro-temporal patterns. However, previous formulations do not model prominent output structure. In this paper, we propose an alternative training target that is explicitly related to mask structure. We first learn a compositional model of the square-root ideal ratio mask that is closely related to the Wiener filter. Instead of directly estimating the ideal mask values, we learn to predict the weights for resulting mask-level spectro-temporal bases, which are then used to generate the estimated masks. In other words, the discriminative model is used to predict the parameters of a generative model of the target of interest. Experimental results show consistent improvements in low SNR conditions by adopting the new training target. Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 2 |
| 2014 | A two-stage approach for improving the perceptual quality of separated speechabstractBinary time-frequency masking and model-based nonnegative matrix factorization (NMF) are two common approaches to speech separation. However, binary masking often suffers from poor perceptual quality, while NMF typically requires pretrained models for both speech and noise and frequently does not perform well. In this paper we examine whether a single or two-stage approach should be used for performing separation. We propose a two-stage algorithm that uses a soft mask in the first stage for separation, and NMF in the second stage for improving perceptual quality where only a speech model needs to be trained. We show that the proposed two-stage approach achieves higher objective perceptual quality and intelligibility compared to related single-stage methods. Donald S. Williamson, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2014 | Robust speaker identification in noisy and reverberant conditionsabstractRobustness of speaker recognition systems is crucial for real-world applications, which typically contain both additive noise and room reverberation. However, the combined effects of additive noise and convolutive reverberation have been rarely studied in speaker identification (SID). This paper addresses this issue in two phases. We first remove background noise through binary masking using a deep neural network classifier. Then we perform robust SID with speaker models trained in selected reverberant conditions, using bounded marginalization and direct masking. Evaluation results show that the proposed system substantially improves SID performance over related systems in a wide range of reverberation time and signal-to-noise ratios. Xiaojia Zhao, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2014 | Binaural deep neural network classification for reverberant speech segregationabstractWhile human listening is robust in complex auditory scenes, current speech segregation algorithms do not perform well in noisy and reverberant environments. This paper addresses the robustness in binaural speech segregation by employing binary classification based on deep neural networks (DNNs). We systematically examine DNN generalization to untrained configurations. Evaluations and comparisons show that DNN based binaural classification produces superior segregation performance in a variety of multisource and reverberant conditions. DeLiang Wang, Runsheng Liu |
INTERSPEECH | 2 |
| 2014 | Boosted deep neural networks and multi-resolution cochleagram features for voice activity detectionabstractVoice activity detection (VAD) is an important frontend of many speech processing systems. In this paper, we describe a new VAD algorithm based on boosted deep neural networks (bDNNs). The proposed algorithm first generates multiple base predictions for a single frame from only one DNN and then aggregates the base predictions for a better prediction of the frame. Moreover, we employ a new acoustic feature, multi-resolution cochleagram (MRCG), that concatenates the cochleagram features at multiple spectrotemporal resolutions and shows superior speech separation results over many acoustic features. Experimental results show that bDNN-based VAD with the MRCG feature outperforms state-of-the-art VADs by a considerable margin. Xiao-Lei Zhang 0001, DeLiang Wang |
INTERSPEECH | 2 |
| 2014 | A feature study for classification-based speech separation at low signal-to-noise ratiosabstractSpeech separation can be formulated as a classification problem. In classification-based speech separation, supervised learning is employed to classify time-frequency units as either speech-dominant or noise-dominant. In very low signal-to-noise ratio (SNR) conditions, acoustic features extracted from a mixture are crucial for correct classification. In this study, we systematically evaluate a range of promising features for classification-based separation using six nonstationary noises at the low SNR level of -5 dB, which is chosen with the goal of improving human speech intelligibility in mind. In addition, we propose a new feature called multi-resolution cochleagram (MRCG). The new feature is constructed by combining four cochleagrams at different spectrotemporal resolutions in order to capture both the local and contextual information. Experimental results show that MRCG gives the best classification results among all evaluated features. In addition, our results indicate that auto-regressive moving average (ARMA) filtering, a post-processing technique for improving automatic speech recognition features, also improves many acoustic features for speech separation. Jitong Chen, Yuxuan Wang 0002, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Neural network based pitch tracking in very noisy speechabstractPitch determination is a fundamental problem in speech processing, which has been studied for decades. However, it is challenging to determinate pitch in strong noise because the harmonic structure is corrupted. In this paper, we estimate pitch using supervised learning, where the probabilistic pitch states are directly learned from noisy speech data. We investigate two alternative neural networks modeling pitch state distribution given observations. The first one is a feedforward deep neural network (DNN), which is trained on static frame-level acoustic features. The second one is a recurrent deep neural network (RNN) which is trained on sequential frame-level features and capable of learning temporal dynamics. Both DNNs and RNNs produce accurate probabilistic outputs of pitch states, which are then connected into pitch contours by Viterbi decoding. Our systematic evaluation shows that the proposed pitch tracking algorithms are robust to different noise conditions and can even be applied to reverberant speech. The proposed approach also significantly outperforms other state-of-the-art pitch tracking algorithms. DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Binaural classification for reverberant speech segregation using deep neural networksabstractSpeech signal degradation in real environments mainly results from room reverberation and concurrent noise. While human listening is robust in complex auditory scenes, current speech segregation algorithms do not perform well in noisy and reverberant environments. We treat the binaural segregation problem as binary classification, and employ deep neural networks (DNNs) for the classification task. The binaural features of the interaural time difference and interaural level difference are used as the main auditory features for classification. The monaural feature of gammatone frequency cepstral coefficients is also used to improve classification performance, especially when interference and target speech are collocated or very close to one another. We systematically examine DNN generalization to untrained spatial configurations. Evaluations and comparisons show that DNN-based binaural classification produces superior segregation performance in a variety of multisource and reverberant conditions. DeLiang Wang, Runsheng Liu, Zhenming Feng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Investigation of Speech Separation as a Front-End for Noise Robust Speech RecognitionabstractRecently, supervised classification has been shown to work well for the task of speech separation. We perform an in-depth evaluation of such techniques as a front-end for noise-robust automatic speech recognition (ASR). The proposed separation front-end consists of two stages. The first stage removes additive noise via time-frequency masking. The second stage addresses channel mismatch and the distortions introduced by the first stage; a non-linear function is learned that maps the masked spectral features to their clean counterpart. Results show that the proposed front-end substantially improves ASR performance when the acoustic models are trained in clean conditions. We also propose a diagonal feature discriminant linear regression (dFDLR) adaptation that can be performed on a per-utterance basis for ASR systems employing deep neural networks and HMM. Results show that dFDLR consistently improves performance in all test conditions. Surprisingly, the best average results are obtained when dFDLR is applied to models trained using noisy log-Mel spectral features from the multi-condition training set. With no channel mismatch, the best results are obtained when the proposed speech separation front-end is used along with multi-condition training using log-Mel features followed by dFDLR adaptation. Both these results are among the best on the Aurora-4 dataset. Arun Narayanan, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | On training targets for supervised speech separationabstractFormulation of speech separation as a supervised learning problem has shown considerable promise. In its simplest form, a supervised learning algorithm, typically a deep neural network, is trained to learn a mapping from noisy features to a time-frequency representation of the target of interest. Traditionally, the ideal binary mask (IBM) is used as the target because of its simplicity and large speech intelligibility gains. The supervised learning framework, however, is not restricted to the use of binary targets. In this study, we evaluate and compare separation results by using different training targets, including the IBM, the target binary mask, the ideal ratio mask (IRM), the short-time Fourier transform spectral magnitude and its corresponding mask (FFT-MASK), and the Gammatone frequency power spectrum. Our results in various test conditions reveal that the two ratio mask targets, the IRM and the FFT-MASK, outperform the other targets in terms of objective intelligibility and quality metrics. In addition, we find that masking based targets, in general, are significantly better than spectral envelope based targets. We also present comparisons with recent methods in non-negative matrix factorization and speech enhancement, which show clear performance advantages of supervised speech separation. Yuxuan Wang 0002, Arun Narayanan, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Robust Speaker Identification in Noisy and Reverberant ConditionsabstractRobustness of speaker recognition systems is crucial for real-world applications, which typically contain both additive noise and room reverberation. However, the combined effects of additive noise and convolutive reverberation have been rarely studied in speaker identification (SID). This paper addresses this issue in two phases. We first remove background noise through binary masking using a deep neural network classifier. Then we perform robust SID with speaker models trained in selected reverberant conditions, on the basis of bounded marginalization and direct masking. Evaluation results show that the proposed system substantially improves SID performance over related systems in a wide range of reverberation time and signal-to-noise ratios. Xiaojia Zhao, Yuxuan Wang 0002, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Remote Sensing Image Segmentation by Combining Spectral and Texture FeaturesabstractWe present a new method for remote sensing image segmentation, which utilizes both spectral and texture information. Linear filters are used to provide enhanced spatial patterns. For each pixel location, we compute combined spectral and texture features using local spectral histograms, which concatenate local histograms of all input bands. We regard each feature as a linear combination of several representative features, each of which corresponds to a segment. Segmentation is given by estimating combination weights, which indicate segment ownership of pixels. We present segmentation solutions where representative features are either known or unknown. We also show that feature dimensions can be greatly reduced via subspace projection. The scale issue is investigated, and an algorithm is presented to automatically select proper scales, which does not require segmentation at multiple-scale levels. Experimental results demonstrate the promise of the proposed method. Jiangye Yuan, DeLiang Wang, Rongxing Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2013 | Learning invariant features for speech separationabstractRecent studies on speech separation show that the ideal binary mask (IBM) substantially improves speech intelligibility in noise. Supervised learning can be used to effectively estimate the IBM. However, supervised learning has trouble dealing with the situations where the probabilistic properties of the training data and the test data do not match, resulting in a challenging issue of generalization whereby the system trained under particular noise conditions may not generalize to new noise conditions. We propose to use a novel metric learning method to learn invariant speech features in the kernel space. As the learned features encode speech-related information that is robust to different noise types, the system is expected to generalize to unseen noise conditions. Evaluations show the advantage of the proposed approach over other speech separation systems. DeLiang Wang |
ICASSP | 2 |
| 2013 | Coupling binary masking and robust ASRabstractWe present a novel framework for performing speech separation and robust automatic speech recognition (ASR) in a unified fashion. Separation is performed by estimating the ideal binary mask (IBM), which identifies speech dominant and noise dominant units in a time-frequency (T-F) representation of the noisy signal. ASR is performed on extracted cepstral features after binary masking. Previous systems perform these steps in a sequential fashion - separation followed by recognition. The proposed framework, which we call bidirectional speech decoding (BSD), unifies these two stages. It does this by using multiple IBM estimators each of which is designed specifically for a back-end acoustic phonetic unit (BPU) of the recognizer. The standard ASR decoder is modified to use these IBM estimators to obtain BPU-specific cepstra during likelihood calculation. On the Aurora-4 robust ASR task, the proposed framework obtains a relative improvement of 17% in word error rate over the noisy baseline. It also obtains significant improvements in the quality of the estimated IBM. Arun Narayanan, DeLiang Wang |
ICASSP | 2 |
| 2013 | Ideal ratio mask estimation using deep neural networks for robust speech recognitionabstractWe propose a feature enhancement algorithm to improve robust automatic speech recognition (ASR). The algorithm estimates a smoothed ideal ratio mask (IRM) in the Mel frequency domain using deep neural networks and a set of time-frequency unit level features that has previously been used to estimate the ideal binary mask. The estimated IRM is used to filter out noise from a noisy Mel spectrogram before performing cepstral feature extraction for ASR. On the noisy subset of the Aurora-4 robust ASR corpus, the proposed enhancement obtains a relative improvement of over 38% in terms of word error rates using ASR models trained in clean conditions, and an improvement of over 14% when the models are trained using the multi-condition training data. In terms of instantaneous SNR estimation performance, the proposed system obtains a mean absolute error of less than 4 dB in most frequency channels. Arun Narayanan, DeLiang Wang |
ICASSP | 2 |
| 2013 | Feature denoising for speech separation in unknown noisy environmentsabstractSpeech separation has been recently formulated as a classification problem. Classification as a form of supervised learning usually performs well on background noises when parts of them are seen in the training set. However, the performance can be significantly worse when generalizing to completely unseen noises. In this study, we present a method that alleviates the generalization issue by attempting to denoise acoustic features before training and testing. We show that a standard multilayer perceptron with proper regularization performs well on this task. Experimental results indicate that the resulting separation system performs significantly better in a variety of unknown noises in low SNR conditions. In a negative SNR condition, we also show that the proposed system produces more intelligible speech according to two recently proposed objective speech intelligibility measures. Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 2 |
| 2013 | A sparse representation approach for perceptual quality improvement of separated speechabstractSpeech separation based on time-frequency masking has been shown to improve intelligibility of speech signals corrupted by noise. A perceived weakness of binary masking is the quality of separated speech. In this paper, an approach for improving the perceptual quality of separated speech from binary masking is proposed. Our approach consists of two stages, where a binary mask is generated in the first stage that effectively performs speech separation. In the second stage, a sparse-representation approach is used to represent the separated signal by a linear combination of Short-time Fourier Transform (STFT) magnitudes that are generated from a clean speech dictionary. Overlap-and-add synthesis is then used to generate an estimate of the speech signal. The performance of the proposed approach is evaluated with the Perceptual Evaluation of Speech Quality (PESQ), which is a standard objective speech quality measure. The proposed algorithm offers considerable improvements in speech quality over binary-masked noisy speech and other reconstruction approaches. Donald S. Williamson, Yuxuan Wang 0002, DeLiang Wang |
ICASSP | 3 |
| 2013 | Analyzing noise robustness of MFCC and GFCC features in speaker identificationabstractAutomatic speaker recognition can achieve a high level of performance in matched training and testing conditions. However, such performance drops significantly in mismatched noisy conditions. Recent research indicates that a new speaker feature, gammatone frequency cepstral coefficients (GFCC), exhibits superior noise robustness to commonly used mel-frequency cepstral coefficients (MFCC). To gain a deep understanding of the intrinsic robustness of GFCC relative to MFCC, we design speaker identification experiments to systematically analyze their differences and similarities. This study reveals that the nonlinear rectification accounts for the noise robustness differences primarily. Moreover, this study suggests how to enhance MFCC robustness, and further improve GFCC robustness by adopting a different time-frequency representation. Xiaojia Zhao, DeLiang Wang |
ICASSP | 2 |
| 2013 | Special issue on advanced theory and methodology in intelligent computing: Selected papers from the Seventh International Conference on Intelligent Computing (ICIC 2011)
De-Shuang Huang, DeLiang Wang |
Neurocomputing | 2 |
| 2013 | Towards Generalizing Classification Based Speech SeparationabstractAbstract—Monaural speech separation is a well-recognized challenge. Recent studies utilize supervised classification methods to estimate the ideal binary mask (IBM) to address the problem. In a supervised learning framework, the issue of generalization to conditions different from those in training is very important. This paper presents methods that require only a small training corpus and can generalize to unseen conditions. The system utilizes support vector machines to learn classification cues and then employs a rethresholding technique to estimate the IBM. A distribution fitting method is used to generalize to unseen signal-to-noise ratio conditions and voice activity detection based adaptationisusedtogeneralizetounseen noise conditions. Systematic evaluation and comparison show that the proposed approach produces high quality IBM estimates under unseen conditions. Index Terms—Generalization, rethresholding, speech separation, support vector machine (SVM). I. DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | A Direct Masking Approach to Robust ASRabstractRecently, much work has been devoted to the computation of binary masks for speech segregation. Conventional wisdom in the field of ASR holds that these binary masks cannot be used directly; the missing energy significantly affects the calculation of the cepstral features commonly used in ASR. We show that this commonly held belief may be a misconception; we demonstrate the effectiveness of directly using the masked data on both a small and large vocabulary dataset. In fact, this approach, which we term the direct masking approach, performs comparably to two previously proposed missing feature techniques. We also investigate the reasons why other researchers may have not come to this conclusion; variance normalization of the features is a significant factor in performance. This work suggests a much better baseline than unenhanced speech for future work in missing feature ASR. William Hartmann, Arun Narayanan, Eric Fosler-Lussier, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 4 |
| 2013 | An Unsupervised Approach to Cochannel Speech SeparationabstractCochannel (two-talker) speech separation is predominantly addressed using pretrained speaker dependent models. In this paper, we propose an unsupervised approach to separating cochannel speech. Our approach follows the two main stages of computational auditory scene analysis: segmentation and grouping. For voiced speech segregation, the proposed system utilizes a tandem algorithm for simultaneous grouping and then unsupervised clustering for sequential grouping. The clustering is performed by a search to maximize the ratio of between- and within-group speaker distances while penalizing within-group concurrent pitches. To segregate unvoiced speech, we first produce unvoiced speech segments based on onset/offset analysis. The segments are grouped using the complementary binary masks of segregated voiced speech. Despite its simplicity, our approach produces significant SNR improvements across a range of input SNR. The proposed system yields competitive performance in comparison to other speaker-independent and model-based methods. DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Exploring Monaural Features for Classification-Based Speech SegregationabstractMonaural speech segregation has been a very challenging problem for decades. By casting speech segregation as a binary classification problem, recent advances have been made in computational auditory scene analysis on segregation of both voiced and unvoiced speech. So far, pitch and amplitude modulation spectrogram have been used as two main kinds of time-frequency (T-F) unit level features in classification. In this paper, we expand T-F unit features to include gammatone frequency cepstral coefficients (GFCC), mel-frequency cepstral coefficients, relative spectral transform (RASTA) and perceptual linear prediction (PLP). Comprehensive comparisons are performed in order to identify effective features for classification-based speech segregation. Our experiments in matched and unmatched test conditions show that these newly included features significantly improve speech segregation performance. Specifically, GFCC and RASTA-PLP are the best single features in matched-noise and unmatched-noise test conditions, respectively. We also find that pitch-based features are crucial for good generalization to unseen environments. To further explore complementarity in terms of discriminative power, we propose to use a group Lasso approach to select complementary features in a principled way. The final combined feature set yields promising results in both matched and unmatched test conditions. Yuxuan Wang 0002, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | Towards Scaling Up Classification-Based Speech SeparationabstractFormulating speech separation as a binary classification problem has been shown to be effective. While good separation performance is achieved in matched test conditions using kernel support vector machines (SVMs), separation in unmatched conditions involving new speakers and environments remains a big challenge. A simple yet effective method to cope with the mismatch is to include many different acoustic conditions into the training set. However, large-scale training is almost intractable for kernel machines due to computational complexity. To enable training on relatively large datasets, we propose to learn more linearly separable and discriminative features from raw acoustic features and train linear SVMs, which are much easier and faster to train than kernel SVMs. For feature learning, we employ standard pre-trained deep neural networks (DNNs). The proposed DNN-SVM system is trained on a variety of acoustic conditions within a reasonable amount of time. Experiments on various test mixtures demonstrate good generalization to unseen speakers and background noises. Yuxuan Wang 0002, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Binaural Detection, Localization, and Segregation in Reverberant Environments Based on Joint Pitch and Azimuth CuesabstractWe propose an approach to binaural detection, localization and segregation of speech based on pitch and azimuth cues. We formulate the problem as a search through a multisource state space across time, where each multisource state encodes the number of active sources, and the azimuth and pitch of each active source. A set of multilayer perceptrons are trained to assign time-frequency units to one of the active sources in each multisource state based jointly on observed pitch and azimuth cues. We develop a novel hidden Markov model framework to estimate the most probable path through the multisource state space. An estimated state path encodes a solution to the detection, localization, pitch estimation and simultaneous organization problems. Segregation is then achieved with an azimuth-based sequential organization stage. We demonstrate that the proposed framework improves segregation relative to several two-microphone comparison systems that are based solely on azimuth cues. Performance gains are consistent across a variety of reverberant conditions. John Woodruff, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | On generalization of classification based speech separationabstractMonaural speech separation is a very challenging problem. Recent studies utilize supervised learning methods to estimate the ideal binary mask (IBM) to solve the problem. In a supervised learning framework, the issue of generalization to conditions different from those used in training is paramount. This paper describes methods that require only a small training corpus but can generalize to unseen conditions. The system utilizes support vector machines to learn classification cues and then employs a rethresholding method to estimate the IBM. A distribution fitting method is used to address unseen signal-to-noise ratio conditions and an iterative voice activity detection is used to address unseen noise conditions. Systematic evaluations show that the proposed approach produces high quality IBM estimates under unseen conditions. DeLiang Wang |
ICASSP | 2 |
| 2012 | SVM-based separation of unvoiced-voiced speech in cochannel conditionsabstractUnvoiced-voiced portions of cochannel speech contain considerable amounts of both voiced and unvoiced speech and play a significant role in separation. Motivated by recent developments in separation of speech from nonspeech noise, we propose a classification-based approach for unvoiced-voiced speech separation. A new feature set consisting of pitch-based features and gammatone frequency cepstral coefficients is proposed to represent the characteristics of a time-frequency unit. The cepstral features do not rely on pitch and are thus more robust than the pitch-based features to pitch estimation errors. Speaker-independent support vector machines are trained for classification. Results based on the TIMIT corpus show that the proposed algorithm significantly improves unvoiced speech segregation compared to a recent algorithm. DeLiang Wang |
ICASSP | 2 |
| 2012 | Binaural speech segregation based on pitch and azimuth trackingabstractWe propose an approach to binaural speech segregation in reverberation based on pitch and azimuth cues. These cues are integrated within a statistical tracking framework to estimate up to two concurrent pitch frequencies and three concurrent azimuth angles. The tracking framework implicitly estimates binary time-frequency masks by solving a data association problem, thereby performing speech segregation. Experimental results show that the proposed approach compares favorably to existing two-microphone systems in spite of less prior information. The benefit of the proposed approach is most pronounced in conditions with substantial reverberation or for closely spaced sources. John Woodruff, DeLiang Wang |
ICASSP | 2 |
| 2012 | On the Role of Binary Mask Pattern in Automatic Speech RecognitionabstractProcessing noisy signals using the ideal binary mask has been shown to improve automatic speech recognition (ASR) performance. In this paper, we present the first study that investigates the role of mask patterns in ASR under varying signalto-noise ratios (SNR), noise conditions and mask definitions. Binary masks are typically computed either by comparing the local SNR within a time-frequency unit of a mixture signal with a threshold termed the local criterion (LC), or by comparing the local target energy with the long-term average energy of speech. Results show that: (i) Akin to human speech recognition, binary masking can significantly improve ASR even when the mixture SNR is as low as -60 dB. (ii) The difference between the LC and the mixture SNR is more correlated to the recognition accuracy than LC. (iii) The performance profiles in ASR are qualitatively similar to those obtained for human speech recognition. (iv) The LC at which the peak performance is obtained is lower than 0 dB, which is the optimal threshold as far as the SNR gain of processed signals is concerned. This indicates that maximizing SNR gain may not be the optimal criterion to improve either human or machine recognition of noisy speech. Arun Narayanan, DeLiang Wang |
INTERSPEECH | 2 |
| 2012 | Acoustic Features for Classification Based Speech Separation
Yuxuan Wang 0002, DeLiang Wang |
INTERSPEECH | 3 |
| 2012 | Boosting Classification Based Speech Separation Using Temporal DynamicsabstractSignificant advances in speech separation have been made by formulating it as a classification problem, where the desired output is the ideal binary mask (IBM). Previous work does not explicitly model the correlation between neighboring time-frequency units and standard binary classifiers are used. As one of the most important characteristics of speech signal is its temporal dynamics, the IBM contains highly structured, instead of, random patterns. In this study, we incorporate temporal dynamics into classification by employing structured output learning. In particular, we use linear-chain structured perceptrons to account for the interactions of neighboring labels in time. However, the performance of structured perceptrons largely depends on the linear separability of features. To address this problem, we employ pretrained deep neural networks to automatically learn effective feature functions for structured perceptrons. The experiments show that the proposed system significantly outperforms previous IBM estimation systems. Index Terms: Monaural speech separation, temporal dynamics, structured perceptron, deep neural networks Yuxuan Wang 0002, DeLiang Wang |
INTERSPEECH | 2 |
| 2012 | Cocktail Party Processing via Structured PredictionabstractWhile human listeners excel at selectively attending to a conversation in a cocktail party, machine performance is still far inferior by comparison. We show that the cocktail party problem, or the speech separation problem, can be effectively approached via structured prediction. To account for temporal dynamics in speech, we employ conditional random fields (CRFs) to classify speech dominance within each time-frequency unit for a sound mixture. To capture complex, nonlinear relationship between input and output, both state and transition feature functions in CRFs are learned by deep neural networks. The formulation of the problem as classification allows us to directly optimize a measure that is well correlated with human speech intelligibility. The proposed system substantially outperforms existing ones in a variety of noises. Yuxuan Wang 0002, DeLiang Wang |
NIPS | 2 |
| 2012 | Expedited review process
Kenji Doya, John G. Taylor, DeLiang Wang |
Neural Networks | 3 |
| 2012 | Loss of a Co-Editor-in-Chief and friend
Kenji Doya, DeLiang Wang |
Neural Networks | 2 |
| 2012 | Image segmentation using local spectral histograms and linear regression
Jiangye Yuan, DeLiang Wang, Rongxing Li |
Pattern Recognit. Lett. | 2 |
| 2012 | A Tandem Algorithm for Singing Pitch Extraction and Voice Separation From Music AccompanimentabstractSinging pitch estimation and singing voice separation are challenging due to the presence of music accompaniments that are often nonstationary and harmonic. Inspired by computational auditory scene analysis (CASA), this paper investigates a tandem algorithm that estimates the singing pitch and separates the singing voice jointly and iteratively. Rough pitches are first estimated and then used to separate the target singer by considering harmonicity and temporal continuity. The separated singing voice and estimated pitches are used to improve each other iteratively. To enhance the performance of the tandem algorithm for dealing with musical recordings, we propose a trend estimation algorithm to detect the pitch ranges of a singing voice in each time frame. The detected trend substantially reduces the difficulty of singing pitch detection by removing a large number of wrong pitch candidates either produced by musical instruments or the overtones of the singing voice. Systematic evaluation shows that the tandem algorithm outperforms previous systems for pitch extraction and singing voice separation. Chao-Ling Hsu, DeLiang Wang, Jyh-Shing Roger Jang |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A CASA-Based System for Long-Term SNR EstimationabstractWe present a system for robust signal-to-noise ratio (SNR) estimation based on computational auditory scene analysis (CASA). The proposed algorithm uses an estimate of the ideal binary mask to segregate a time-frequency representation of the noisy signal into speech dominated and noise dominated regions. Energy within each of these regions is summated to derive the filtered global SNR. An SNR transform is introduced to convert the estimated filtered SNR to the true broadband SNR of the noisy signal. The algorithm is further extended to estimate subband SNRs. Evaluations are done using the TIMIT speech corpus and the NOISEX92 noise database. Results indicate that both global and subband SNR estimates are superior to those of existing methods, especially at low SNR conditions. Arun Narayanan, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Binaural Localization of Multiple Sources in Reverberant and Noisy EnvironmentsabstractSound source localization from a binaural input is a challenging problem, particularly when multiple sources are active simultaneously and reverberation or background noise are present. In this work, we investigate a multi-source localization framework in which monaural source segregation is used as a mechanism to increase the robustness of azimuth estimates from a binaural input. We demonstrate performance improvement relative to binaural only methods assuming a known number of spatially stationary sources. We also propose a flexible azimuth-dependent model of binaural features that independently captures characteristics of the binaural setup and environmental conditions, allowing for adaptation to new environments or calibration to an unseen binaural setup. Results with both simulated and recorded impulse responses show that robust performance can be achieved with limited prior training. John Woodruff, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | CASA-Based Robust Speaker IdentificationabstractConventional speaker recognition systems perform poorly under noisy conditions. Inspired by auditory perception, computational auditory scene analysis (CASA) typically segregates speech by producing a binary time-frequency mask. We investigate CASA for robust speaker identification. We first introduce a novel speaker feature, gammatone frequency cepstral coefficient (GFCC), based on an auditory periphery model, and show that this feature captures speaker characteristics and performs substantially better than conventional speaker features under noisy conditions. To deal with noisy speech, we apply CASA separation and then either reconstruct or marginalize corrupted components indicated by a CASA mask. We find that both reconstruction and marginalization are effective. We further combine the two methods into a single system based on their complementary advantages, and this system achieves significant performance improvements over related systems under a wide range of signal-to-noise ratios. Xiaojia Zhao, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | An SVM based classification approach to speech separationabstractMonaural speech separation is a very challenging task. CASA-based systems utilize acoustic features to produce a time-frequency (T-F) mask. In this study, we propose a classification approach to monaural separation problem. Our feature set consists of pitch-based features and amplitude modulation spectrum features, which can discriminate both voiced and unvoiced speech from nonspeech interference. We employ support vector machines (SVMs) followed by a re-thresholding method to classify each T-F unit as either target-dominated or interference-dominated. An auditory segmentation stage is then utilized to improve SVM-generated results. Systematic evaluations show that our approach produces high quality binary masks and outperforms a previous system in terms of classification accuracy. DeLiang Wang |
ICASSP | 2 |
| 2011 | A trend estimation algorithm for singing pitch detection in musical recordingsabstractDetecting pitch values for singing voice in the presence of music accompaniment is challenging but useful for many applications. We propose a trend estimation algorithm to detect the pitch ranges of a singing voice in each time frame. The detected trend substantially reduces the difficulty of singing pitch detection by reducing a large number of wrong pitch candidates either produced by musical instruments or the overtones of the singing voice. The proposed algorithm can be applied to improve the performance of singing pitch detection. Quantitative evaluations show that proposed trend estimation improves an existing algorithm significantly. The results from the MIREX 2010 competition show that our system achieves the best overall raw-pitch accuracy for vocal songs. Chao-Ling Hsu, DeLiang Wang, Jyh-Shing Roger Jang |
ICASSP | 2 |
| 2011 | An approach to sequential grouping in cochannel speechabstractModel-based methods for sequential organization in cochannel speech require pretrained speaker models and often prior knowledge of participating speakers. We propose an unsupervised approach to sequential organization of cochannel speech. Based on cepstral features, we first cluster voiced speech into two speaker groups by maximizing the ratio of between- and within-group distances penalized by within-group concurrent pitches. To group unvoiced speech, we employ an onset/offset based analysis to generate time-frequency segments. Unvoiced segments are then labeled by the complementary portions of segregated voiced speech. Our method does not require any pretrained model and is computationally simple. Evaluations and comparisons show that the proposed method outperforms a model-based method in terms of speech segregation. DeLiang Wang |
ICASSP | 2 |
| 2011 | On the use of ideal binary masks for improving phonetic classificationabstractIdeal binary masks are binary patterns that encode the masking characteristics of speech in noise. Recent evidence in speech perception suggests that such binary patterns provide sufficient information for human speech recognition. Motivated by these findings, we propose to use ideal binary masks to improve phonetic modeling. We show that by combining the outputs of classifiers trained on the traditional MFCC features and this novel speech pattern, statistically significant improvements over the baseline MFCC based classifier can be achieved for the task of phonetic classification. Using the combined classifiers, we achieve an error rate of 19.5% on the TIMIT phonetic classification task using multilayer perceptrons as the underlying classifier. Arun Narayanan, DeLiang Wang |
ICASSP | 2 |
| 2011 | Robust speech recognition using multiple prior models for speech reconstructionabstractPrior models of speech have been used in robust automatic speech recognition to enhance noisy speech. Typically, a single prior model is trained by pooling the entire training data. In this paper we propose to train multiple prior models of speech instead of a single prior model. The prior models can be trained based on distinct characteristics of speech. In this study, they are trained based on voicing characteristics. The trained prior models are then used to reconstruct noisy speech. Significant improvements are obtained on the Aurora-4 robust speech recognition task when multiple priors are used; in conjunction with an uncertainty transform technique, multiple priors yield a 13.7% absolute improvement in the average word error rate over directly recognizing noisy speech. Arun Narayanan, Xiaojia Zhao, DeLiang Wang, Eric Fosler-Lussier |
ICASSP | 3 |
| 2011 | Directionality-based speech enhancement for hearing aidsabstractIn this work we describe methods for using the directionality of sound energy as a criterion to estimate single- and multichannel linear filters for suppression of diffuse noise and reverberation in a hearing aid application. We compare conservative strategies where direction of arrival is unknown, and more aggressive strategies where the proposed methods can be used to derive a fast acting post-filter for the output of a beamformer. We show that in situations where a target of interest is near to the listener while interfering sources are more distant, simple features that capture the directionality of sound energy can be used to attenuate significant undesired signal energy and can be more effective than a strategy based on noise-floor tracking. John Woodruff, DeLiang Wang |
ICASSP | 2 |
| 2011 | Robust speaker identification using a CASA front-endabstractSpeaker recognition remains a challenging task under noisy conditions. Inspired by auditory perception, computational auditory scene analysis (CASA) typically segregates speech by producing a binary time-frequency mask. We first show that a recently introduced speaker feature, Gammatone Frequency Cepstral Coefficient, performs substantially better than conventional speaker features under noisy conditions. To deal with noisy speech, we apply CASA separation and then either reconstruct or marginalize corrupted components indicated by the CASA mask. Both methods are effective. We further combine them into a single system depending on the detected signal to noise ratio (SNR). This system achieves significant performance improvements over related systems under a wide range of SNR conditions. Xiaojia Zhao, DeLiang Wang |
ICASSP | 3 |
| 2011 | Image segmentation based on local spectral histograms and linear regressionabstractWe present a novel method for segmenting images with texture and nontexture regions. Local spectral histograms are feature vectors consisting of histograms of chosen filter responses, which capture both texture and nontexture information. Based on the observation that the local spectral histogram of a pixel location can be approximated through a linear combination of the representative features weighted by the area coverage of each feature, we formulate the segmentation problem as a multivariate linear regression, where the solution is obtained by least squares estimation. Moreover, we propose an algorithm to automatically identify representative features corresponding to different homogeneous regions, and show that the number of representative features can be determined by examining the effective rank of a feature matrix. We present segmentation results on different types of images, and our comparison with another spectral histogram based method shows that the proposed method gives more accurate results. Jiangye Yuan, DeLiang Wang, Rongxing Li |
IJCNN | 2 |
| 2011 | An excellent year and a transition
Kenji Doya, Stephen Grossberg, John G. Taylor, DeLiang Wang |
Neural Networks | 4 |
| 2011 | Selecting salient objects in real scenes: An oscillatory correlation model
Marcos G. Quiles, DeLiang Wang, Liang Zhao 0001, Roseli A. Francelin Romero, De-Shuang Huang |
Neural Networks | 2 |
| 2011 | A multistage approach to blind separation of convolutive speech mixtures
Tariqullah Jan, Wenwu Wang 0001, DeLiang Wang |
Speech Commun. | 3 |
| 2011 | Unvoiced Speech Segregation From Nonspeech Interference via CASA and Spectral SubtractionabstractWhile a lot of effort has been made in computational auditory scene analysis to segregate voiced speech from monaural mixtures, unvoiced speech segregation has not received much attention. Unvoiced speech is highly susceptible to interference due to its relatively weak energy and lack of harmonic structure, and hence makes its segregation extremely difficult. This paper proposes a new approach to segregation of unvoiced speech from nonspeech interference. The proposed system first removes estimated voiced speech, and the periodic part of interference based on cross-channel correlation. The resultant interference becomes more stationary and we estimate the noise energy in unvoiced intervals using segregated speech in neighboring voiced intervals. Then unvoiced speech segregation occurs in two stages: segmentation and grouping. In segmentation, we apply spectral subtraction to generate time-frequency segments in unvoiced intervals. Unvoiced speech segments are subsequently grouped based on frequency characteristics of unvoiced speech using simple thresholding as well as Bayesian classification. The proposed algorithm is computationally efficient, and systematic evaluation and comparison show that our approach considerably improves the performance of unvoiced speech segregation. DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | HMM-Based Multipitch Tracking for Noisy and Reverberant SpeechabstractMultipitch tracking in real environments is critical for speech signal processing. Determining pitch in reverberant and noisy speech is a particularly challenging task. In this paper, we propose a robust algorithm for multipitch tracking in the presence of both background noise and room reverberation. An auditory front-end and a new channel selection method are utilized to extract periodicity features. We derive pitch scores for each pitch state, which estimate the likelihoods of the observed periodicity features given pitch candidates. A hidden Markov model integrates these pitch scores and searches for the best pitch state sequence. Our algorithm can reliably detect single and double pitch contours in noisy and reverberant conditions. Quantitative evaluations show that our approach outperforms existing ones, particularly in reverberant conditions. Zhaozhang Jin, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Reverberant Speech Segregation Based on Multipitch Tracking and ClassificationabstractRoom reverberation creates a major challenge to speech segregation. We propose a computational auditory scene analysis approach to monaural segregation of reverberant voiced speech, which performs multipitch tracking of reverberant mixtures and supervised classification. Speech and nonspeech models are separately trained, and each learns to map from a set of pitch-based features to a grouping cue which encodes the posterior probability of a time-frequency (T-F) unit being dominated by the source with the given pitch estimate. Because interference may be either speech or nonspeech, a likelihood ratio test selects the correct model for labeling corresponding T-F units. Experimental results show that the proposed system performs robustly in different types of interference and various reverberant conditions, and has a significant advantage over existing systems. Zhaozhang Jin, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | LEGION-Based Automatic Road Extraction From Satellite ImageryabstractAn automatic method for road extraction from satellite imagery is presented. The core of the proposed method is locally excitatory globally inhibitory oscillator networks (LEGION). The road extraction task is decomposed into three stages. The first stage is image segmentation by LEGION. In the second stage, the medial axis of each segment is computed, and the medial axis points corresponding to narrow regions are selected. The third is the road grouping stage. Alignment-dependent connections between selected points are established, and LEGION is utilized to group well-aligned points, which represent the extracted roads. Due to the selective gating mechanism of LEGION, different roads in an image are grouped separately. Road extraction results on synthetic and real images are presented. A comparison with other methods shows that the proposed method produces very competitive extraction results. Jiangye Yuan, DeLiang Wang, Bo Wu 0004, Lin Yan 0001, Rongxing Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2010 | A multipitch tracking algorithm for noisy and reverberant speechabstractDetermining multiple pitches in noisy and reverberant speech is an important and challenging task. We propose a robust multipitch tracking algorithm in the presence of both background noise and room reverberation. A new channel selection method is utilized in conjunction with an auditory front-end to extract periodicity features in the time-frequency space. These features are combined to formulate frame level conditional probabilities given each pitch state. A hidden Markov model is then applied to integrate these probabilities and search for the most likely pitch state sequences. The proposed approach can reliably detect up to two simultaneous pitch contours in noisy and reverberant conditions. Quantitative evaluations show that our system significantly outperforms existing ones, particularly in reverberant environments. Zhaozhang Jin, DeLiang Wang |
ICASSP | 2 |
| 2010 | Integrating monaural and binaural analysis for localizing multiple reverberant sound sourcesabstractLocalization of simultaneous sound sources in natural environments with only two microphones is a challenging problem. Reverberation degrades performance of localization based exclusively on directional cues. We present an approach that integrates monaural and binaural analysis to improve localization of multiple speech sources in noisy and reverberant environments. Our approach incorporates pitch-based monaural processing to perform simultaneous organization of voiced speech. We propose a probabilistic framework to jointly perform localization and sequential organization using binaural cues. We evaluate our system on multi-source speech mixtures in the presence of reverberation and diffuse noise and compare it to two localization approaches that do not incorporate monaural cues. Results indicate that our system can accurately localize multiple sources in very challenging conditions. John Woodruff, DeLiang Wang |
ICASSP | 2 |
| 2010 | Unvoiced speech segregation based on CASA and spectral subtractionabstractUnvoiced speech separation is an important and challenging problem that has not received much attention. We propose a CASA based approach to segregate unvoiced speech from nonspeech interference. As unvoiced speech does not contain periodic signals, we first remove the periodic portions of a mixture including voiced speech. With periodic components removed, the remaining interference becomes more stationary. We estimate the noise energy in unvoiced intervals on the basis of segregated voiced speech. Spectral subtraction is employed to extract time-frequency segments in unvoiced intervals, and we group the segments dominated by unvoiced speech by simple thresholding or Bayesian classification. Systematic evaluation and comparison show that the proposed method considerably improves the unvoiced speech segregation performance under various SNR conditions. DeLiang Wang |
INTERSPEECH | 2 |
| 2010 | Unsupervised sequential organization for cochannel speech separationabstractThe problem of sequential organization in the cochannel speech situation has previously been studied using speaker-model based methods. A major limitation of these methods is that they require the availability of pretrained speaker models and prior knowledge (or detection) of participating speakers. We propose an unsupervised clustering approach to cochannel speech sequential organization. Given enhanced cepstral features, we search for the optimal assignment of simultaneous speech streams by maximizing the betweenand within-cluster scatter matrix ratio penalized by concurrent pitches within individual speakers. A genetic algorithm is employed to speed up the search. Our method does not require trained speaker models, and experiments with both ideal and estimated simultaneous streams show the proposed method outperforms a speakermodel based method in both speech segregation and computational efficiency. DeLiang Wang |
INTERSPEECH | 2 |
| 2010 | Combining monaural and binaural evidence for reverberant speech segregationabstractMost existing binaural approaches to speech segregation rely on spatial filtering. In environments with minimal reverberation and when sources are well separated in space, spatial filtering can achieve excellent results. However, in everyday environments performance degrades substantially. To address these limitations, we incorporate monaural analysis within a binaural segregation system. We use monaural cues to perform both local and across frequency grouping of mixture components, allowing for a more robust application of spatial filtering. We propose a novel framework in which we combine monaural grouping evidence and binaural localization evidence in a linear model for the estimation of the ideal binary mask. Results indicate that with appropriately designed features that capture both monaural and binaural evidence, an extremely simple model achieves a signal-to-noise ratio improvement of up to 3.6 dB relative to using spatial filtering alone. Index Terms: Speech segregation, binaural localization, monaural grouping, linear model John Woodruff, Rohit Prabhavalkar, Eric Fosler-Lussier, DeLiang Wang |
INTERSPEECH | 4 |
| 2010 | A computational auditory scene analysis system for speech segregation and robust speech recognition
Soundararajan Srinivasan, Zhaozhang Jin, DeLiang Wang |
Comput. Speech Lang. | 4 |
| 2010 | Robust speech recognition by integrating speech separation and hypothesis testing
Soundararajan Srinivasan, DeLiang Wang |
Speech Commun. | 2 |
| 2010 | A Tandem Algorithm for Pitch Estimation and Voiced Speech SegregationabstractA lot of effort has been made in computational auditory scene analysis (CASA) to segregate speech from monaural mixtures. The performance of current CASA systems on voiced speech segregation is limited by lacking a robust algorithm for pitch estimation. We propose a tandem algorithm that performs pitch estimation of a target utterance and segregation of voiced portions of target speech jointly and iteratively. This algorithm first obtains a rough estimate of target pitch, and then uses this estimate to segregate target speech using harmonicity and temporal continuity. It then improves both pitch estimation and voiced speech segregation iteratively. Novel methods are proposed for performing segregation with a given pitch estimate and pitch determination with given segregation. Systematic evaluation shows that the tandem algorithm extracts a majority of target speech without including much interference, and it performs substantially better than previous systems for either pitch extraction or voiced speech segregation. Guoning Hu, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Sequential Organization of Speech in Reverberant Environments by Integrating Monaural Grouping and Binaural LocalizationabstractExisting binaural approaches to speech segregation place an exclusive burden on cues related to the location of sound sources in space. These approaches can achieve excellent performance in anechoic conditions but degrade rapidly in realistic environments where room reverberation corrupts localization cues. In this paper, we propose to integrate monaural and binaural processing to achieve segregation and localization of voiced speech in reverberant environments. The proposed approach builds on monaural analysis for simultaneous organization, and combines it with a novel method for generation of location-based cues in a probabilistic framework that jointly achieves localization and sequential organization. We compare localization performance to two existing methods, sequential organization performance to a model-based system that uses only monaural cues, and segregation performance to an exclusively binaural system. Results suggest that the proposed framework allows for improved source localization and robust segregation of voiced speech in environments with considerable reverberation. John Woodruff, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Incorporating spectral subtraction and noise type for unvoiced speech segregationabstractUnvoiced speech poses a big challenge to current monaural speech segregation systems. It lacks harmonic structure and is highly susceptible to interference due to its relatively weak energy. This paper describes a new approach to segregate unvoiced speech from nonspeech interference. The system first estimates a voiced binary mask, and then performs unvoiced speech segregation in two stages: segmentation and grouping. In segmentation, time-frequency units labeled as 0 in the voiced binary mask are first used to estimate the noise energy and spectral subtraction is then performed to generate time-frequency segments in unvoiced intervals. Based on the type of noise, unvoiced segments are grouped either by selecting segments consistent with those generated by onset/offset analysis or by Bayesian classification of acoustic-phonetic features. Systematic evaluation and comparison show that the proposed approach improves the performance of unvoiced speech segregation considerably. DeLiang Wang |
ICASSP | 2 |
| 2009 | A multistage approach for blind separation of convolutive speech mixturesabstractIn this paper, we propose a novel algorithm for the separation of convolutive speech mixtures using two-microphone recordings, based on the combination of independent component analysis (ICA) and ideal binary mask (IBM), together with a post-filtering process in the cepstral domain. Essentially, the proposed algorithm consists of three steps. First, a constrained convolutive ICA algorithm is applied to separate the source signals from two-microphone recordings. In the second step, we estimate the IBM by comparing the energy of corresponding time-frequency (T-F) units from the separated sources obtained with the convolutive ICA algorithm. The last step is to reduce musical noise caused typically by T-F masking using cepstral smoothing. The performance of the proposed approach is evaluated based on both reverberant mixtures generated using a simulated room model and real recordings. The proposed algorithm offers considerably higher efficiency, together with improved speech quality while producing similar separation performance as compared with a recent approach. Tariqullah Jan, Wenwu Wang 0001, DeLiang Wang |
ICASSP | 3 |
| 2009 | Learning to maximize signal-to-noise ratio for reverberant speech segregationabstractMonaural speech segregation in reverberant environments is a very difficult problem. We develop a supervised learning approach by proposing an objective function that directly relates to the computational goal of maximizing signal-to-noise ratio. The model trained using this new objective function yields significantly better results for time-frequency unit labeling. In our segregation system, a segmentation and grouping framework is utilized to form reliable segments under reverberant conditions and organize them into streams. Systematic evaluations show very promising results. Zhaozhang Jin, DeLiang Wang |
ICASSP | 2 |
| 2009 | An auditory-based feature for robust speech recognitionabstractA conventional automatic speech recognizer does not perform well in the presence of noise, while human listeners are able to segregate and recognize speech in noisy conditions. We study a novel feature based on an auditory periphery model for robust speech recognition. Specifically, gammatone frequency cepstral coefficients are derived by applying a cepstral analysis on gammatone filterbank responses. Our evaluations show that the proposed feature performs considerably better than conventional acoustic features. We further demonstrate that integrating the proposed feature with a computational auditory scene analysis system yields promising recognition performance. Zhaozhang Jin, DeLiang Wang, Soundararajan Srinivasan |
ICASSP | 3 |
| 2009 | On the role of localization cues in binaural segregation of reverberant speechabstractApproaches to binaural and stereo speech segregation have often assumed that localization information can be used as a primary cue to achieve segregation of a target signal. Results produced by these systems degrade significantly in the presence of room reverberation. In this work, we present an alternative framework to achieve localization of groups of time-frequency units. We show that grouping across time and frequency allows the use of localization as an important cue for sequential grouping of time-frequency objects. We analyze the level of time-frequency grouping needed to achieve accurate object localization and show preliminary binaural segregation results using the proposed framework. Results indicate that both localization and segregation performance can be improved by grouping across time and frequency. John Woodruff, DeLiang Wang |
ICASSP | 2 |
| 2009 | An oscillatory correlation model of object-based attentionabstractAttention is a critical mechanism for visual scene analysis. By means of attention, it is possible to break down the analysis of a complex scene to the analysis of its parts through a selection process. Empirical studies demonstrate that attentional selection is conducted on visual objects as a whole. We present a neurocomputational model of object-based selection in the framework of oscillatory correlation. By segmenting an input scene and integrating the segments with their conspicuity obtained from a saliency map, the model selects salient objects rather than salient locations. The proposed system is composed of three modules: a saliency map providing saliency values of image locations, image segmentation for breaking the input scene into a set of objects, and object selection which allows one of the objects of the scene to be selected at a time. This object selection system has been applied to real images and the simulation results show its effectiveness. Marcos G. Quiles, DeLiang Wang, Liang Zhao 0001, Roseli A. Francelin Romero, De-Shuang Huang |
IJCNN | 2 |
| 2009 | Automatic road extraction from satellite imagery using LEGION networksabstractWe present an automatic method for road extraction from satellite imagery. The core of the proposed method is locally excitatory globally inhibitory oscillator networks (LEGION). We decompose the road extraction task into three stages. The first stage is image segmentation by LEGION. In the second stage, we compute the medial axis of each segment and select the segments with narrow widths. The third is the road grouping stage. With the medial axes, alignment-dependent connections between medial axis points are established and LEGION is utilized to group the well-aligned medial axes, which represent extracted road segments. Due to the selective gating mechanism of LEGION, different roads in an image are grouped separately. Experimental results on synthetic and real images show the effectiveness of this method. Jiangye Yuan, DeLiang Wang, Bo Wu 0004, Lin Yan 0001, Rongxing Li |
IJCNN | 2 |
| 2009 | On the optimality of ideal binary time-frequency masks
DeLiang Wang |
Speech Commun. | 2 |
| 2009 | Sequential organization of speech in computational auditory scene analysis
DeLiang Wang |
Speech Commun. | 2 |
| 2009 | A Supervised Learning Approach to Monaural Segregation of Reverberant SpeechabstractA major source of signal degradation in real environments is room reverberation. Monaural speech segregation in reverberant environments is a particularly challenging problem. Although inverse filtering has been proposed to partially restore the harmonicity of reverberant speech before segregation, this approach is sensitive to specific source/receiver and room configurations. This paper proposes a supervised learning approach to monaural segregation of reverberant voiced speech, which learns to map from a set of pitch-based auditory features to a grouping cue encoding the posterior probability of a time-frequency (T-F) unit being target dominant given observed features. We devise a novel objective function for the learning process, which directly relates to the goal of maximizing signal-to-noise ratio. The models trained using this objective function yield significantly better T-F unit labeling. A segmentation and grouping framework is utilized to form reliable segments under reverberant conditions and organize them into streams. Systematic evaluations show that our approach produces very promising results under various reverberant conditions and generalizes well to new utterances and new speakers. Zhaozhang Jin, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Monaural Musical Sound Separation Based on Pitch and Common Amplitude ModulationabstractMonaural musical sound separation has been extensively studied recently. An important problem in separation of pitched musical sounds is the estimation of time-frequency regions where harmonics overlap. In this paper, we propose a sinusoidal modeling-based separation system that can effectively resolve overlapping harmonics. Our strategy is based on the observations that harmonics of the same source have correlated amplitude envelopes and that the change in phase of a harmonic is related to the instrument's pitch. We use these two observations in a least squares estimation framework for separation of overlapping harmonics. The system directly distributes mixture energy for harmonics that are unobstructed by other sources. Quantitative evaluation of the proposed system is shown when ground truth pitch information is available, when rough pitch estimates are provided in the form of a MIDI score, and finally, when a multi pitch tracking algorithm is used. We also introduce a technique to improve the accuracy of rough pitch estimates. Results show that the proposed system significantly outperforms related monaural musical sound separation systems. John Woodruff, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Musical Sound Separation Using Pitch-Based Labeling and Binary Time-Frequency MaskingabstractMonaural musical sound separation attempts to segregate different instrument lines from single-channel polyphonic music. We propose a system that decomposes an input into time-frequency units using an auditory filterbank and utilizes pitch to label which instrument line each time-frequency unit is assigned to. The system is conceptually simple and computationally efficient. Systematic evaluation shows that, despite its simplicity, the proposed system achieves a competitive level of performance. DeLiang Wang |
ICASSP | 2 |
| 2008 | On the optimality of ideal binary time-frequency masksabstractRecently the concept of ideal binary time-frequency masks has received attention and their optimality in terms of signal- to-noise ratio has been presumed. However the optimality is not rigorously analyzed. In this paper we treat this issue formally and clarify the conditions for ideal binary masks to be optimal. We also experimentally compare the performance of ideal binary masks in terms of signal-to-noise ratio to that of ideal ratio masks on a speech mixture database and a music database. The results show that ideal binary masks are close in performance to ideal ratio masks which are closely related to the Wiener filter, the theoretically optimal linear filter. DeLiang Wang |
ICASSP | 2 |
| 2008 | Robust speaker identification using auditory features and computational auditory scene analysisabstractThe performance of speaker recognition systems drop significantly under noisy conditions. To improve robustness, we have recently proposed novel auditory features and a robust speaker recognition system using a front-end based on computational auditory scene analysis. In this paper, we further study the auditory features by exploring different feature dimensions and incorporating dynamic features. In addition, we evaluate the features and robust recognition in a speaker identification task in a number of noisy conditions. We find that one of the auditory features performs substantially better than a conventional speaker feature. Furthermore, our recognition system achieves significant performance improvements compared with an advanced front-end in a wide range of signal-to-noise conditions. DeLiang Wang |
ICASSP | 2 |
| 2008 | Binaural Tracking of Multiple Moving SourcesabstractThis paper addresses the problem of tracking multiple moving sources using binaural input. We observe that binaural cues are strongly correlated with source locations in time-frequency regions dominated by only one source. Based on this observation, we propose a novel tracking algorithm that integrates probabilities across reliable frequency channels in order to produce a likelihood function in the target space, which describes the azimuths of all active sources at a particular time frame. Finally, a hidden Markov model (HMM) is employed to form continuous tracks and automatically detect the number of active sources across time. Results are presented for up to three moving talkers in anechoic conditions. A comparison shows that our HMM model outperforms a Kalman filter-based approach in tracking active sources across time. Our study represents a first step in addressing auditory scene analysis with moving sound sources. Nicoleta Roman, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Two-Microphone Separation of Speech MixturesabstractSeparation of speech mixtures, often referred to as the cocktail party problem, has been studied for decades. In many source separation tasks, the separation method is limited by the assumption of at least as many sensors as sources. Further, many methods require that the number of signals within the recorded mixtures be known in advance. In many real-world applications, these limitations are too restrictive. We propose a novel method for underdetermined blind source separation using an instantaneous mixing model which assumes closely spaced microphones. Two source separation techniques have been combined, independent component analysis (ICA) and binary time - frequency (T-F) masking. By estimating binary masks from the outputs of an ICA algorithm, it is possible in an iterative way to extract basis speech signals from a convolutive mixture. The basis signals are afterwards improved by grouping similar signals. Using two microphones, we can separate, in principle, an arbitrary number of mixed speech signals. We show separation results for mixtures with as many as seven speech signals under instantaneous conditions. We also show that the proposed method is applicable to segregate speech signals under reverberant conditions, and we compare our proposed method to another state-of-the-art algorithm. The number of source signals is not assumed to be known in advance and it is possible to maintain the extracted signals as stereo signals. Michael Syskind Pedersen, DeLiang Wang, Jan Larsen, Ulrik Kjems |
IEEE Trans. Neural Networks | 2 |
| 2007 | A Supervised Learning Approach to Monaural Segregation of Reverberant SpeechabstractRoom reverberation degrades speech signals and poses a major challenge to current monaural speech segregation systems. Previous research relies on inverse filtering as a front-end for partially restoring the harmonicity of the reverberant signal. We show that the inverse filtering approach is sensitive to different room configurations, hence undesirable in general reverberation conditions. We propose a supervised learning approach to map a set of harmonic features into a pitch based grouping cue for each time-frequency (T-F) unit. We use a speech segregation method to estimate an ideal binary T-F mask which retains the reverberant mixture in a local T-F unit if and only if the energy of target is stronger than interference energy. Results show that our approach improves the segregation performance considerably. Zhaozhang Jin, DeLiang Wang |
ICASSP (4) | 2 |
| 2007 | Pitch Detection in Polyphonic Music using Instrument Tone ModelsabstractWe propose a hidden Markov model (HMM) based system to detect the pitch of an instrument in polyphonic music using an instrument tone model. Our system calculates at every time frame the salience of a pitch hypothesis based on the magnitudes of harmonics associated with the hypothesis. A hypothesis selection method is introduced to choose pitch hypotheses with sufficiently high salience as pitch candidates. Then the system applies an instrument model to evaluate the likelihood of each candidate. The transition probability between successive pitch points is constructed using the prior knowledge of the musical key of the input. Finally an HMM integrates the instrument likelihood and the pitch transition probability. Quantitative evaluation shows the proposed system performs well for different instruments. We also compare a Gaussian mixture model and kernel density estimation for instrument modeling, and find that kernel density estimation gives better overall performance while the Gaussian mixture model is more robust. DeLiang Wang |
ICASSP (2) | 2 |
| 2007 | Incorporating Auditory Feature Uncertainties in Robust Speaker IdentificationabstractConventional speaker recognition systems perform poorly under noisy conditions. Recent research suggests that binary time-frequency (T-F) masks be a promising front-end for robust speaker recognition. In this paper, we propose novel auditory features based on an auditory periphery model, and show that these features capture significant speaker characteristics. Additionally, we estimate uncertainties of the auditory features based on binary T-F masks, and calculate speaker likelihood scores using uncertainty decoding. Our approach achieves substantial performance improvement in a speaker identification task compared with a state-of-the-art robust front-end in a wide range of signal-to-noise conditions. Soundararajan Srinivasan, DeLiang Wang |
ICASSP (4) | 3 |
| 2007 | Exploiting Uncertainties for Binaural Speech RecognitionabstractRecently several algorithms have been proposed to enhance noisy speech by estimating the signal-to-noise ratio (SNR) within a local time-frequency region based on binaural cues of interaural time and intensity differences (ITD and IID). However, the accuracy of the estimated SNR often varies widely across time and frequency, causing uncertainties in the enhanced speech features. We estimate this uncertainty based on statistics of ITD and IID and show that it can be effectively exploited to improve robust speech recognition. Systematic evaluations using the estimated uncertainty show significant improvement in recognition performance compared to the baseline performance. Soundararajan Srinivasan, Nicoleta Roman, DeLiang Wang |
ICASSP (4) | 3 |
| 2007 | Auditory Segmentation Based on Onset and Offset AnalysisabstractA typical auditory scene in a natural environment contains multiple sources. Auditory scene analysis (ASA) is the process in which the auditory system segregates a scene into streams corresponding to different sources. Segmentation is a major stage of ASA by which an auditory scene is decomposed into segments, each containing signal mainly from one source. We propose a system for auditory segmentation by analyzing onsets and offsets of auditory events. The proposed system first detects onsets and offsets, and then generates segments by matching corresponding onset and offset fronts. This is achieved through a multiscale approach. A quantitative measure is suggested for segmentation evaluation. Systematic evaluation shows that most of target speech, including unvoiced speech, is correctly segmented, and target speech and interference are well separated into different segments Guoning Hu, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Separation of Singing Voice From Music Accompaniment for Monaural RecordingsabstractSeparating singing voice from music accompaniment is very useful in many applications, such as lyrics recognition and alignment, singer identification, and music information retrieval. Although speech separation has been extensively studied for decades, singing voice separation has been little investigated. We propose a system to separate singing voice from music accompaniment for monaural recordings. Our system consists of three stages. The singing voice detection stage partitions and classifies an input into vocal and nonvocal portions. For vocal portions, the predominant pitch detection stage detects the pitch of the singing voice and then the separation stage uses the detected pitch to group the time-frequency segments of the singing voice. Quantitative results show that the system performs the separation task successfully DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Transforming Binary Uncertainties for Robust Speech RecognitionabstractRecently, several algorithms have been proposed to enhance noisy speech by estimating a binary mask that can be used to select those time-frequency regions of a noisy speech signal that contain more speech energy than noise energy. This binary mask encodes the uncertainty associated with enhanced speech in the linear spectral domain. The use of the cepstral transformation smears the information from the noise dominant time-frequency regions across all the cepstral features. We propose a supervised approach using regression trees to learn the nonlinear transformation of the uncertainty from the linear spectral domain to the cepstral domain. This uncertainty is used by a decoder that exploits the variance associated with the enhanced cepstral features to improve robust speech recognition. Systematic evaluations on a subset of the Aurora4 task using the estimated uncertainty show substantial improvement over the baseline performance across various noise conditions. Soundararajan Srinivasan, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Speech Recognition in Multisource Reverberant Environments with Binaural InputsabstractWe present a binaural solution to robust speech recognition in multi-source reverberant environments. We employ the notion of an ideal time-frequency binary mask, which selects the target if it is stronger than the interference in a local time-frequency (T-F) unit. Our system estimates this ideal binary mask at the output of a target cancellation module implemented using adaptive filtering. This mask is used in conjunction with a missing-data algorithm to decode the target utterance. A systematic evaluation in terms of automatic speech recognition (ASR) performance shows substantial improvements over the baseline performance and better results over related two-microphone approaches. Nicoleta Roman, Soundararajan Srinivasan, DeLiang Wang |
ICASSP (1) | 3 |
| 2006 | Robust Speaker Recognition Using Binary Time-Frequency MasksabstractConventional speaker recognition systems perform poorly under noisy conditions. In this paper, we evaluate binary time-frequency masks for robust speaker recognition. An ideal binary mask is a priori defined as a binary matrix where 1 indicates that the target is stronger than the interference within the corresponding time-frequency unit and 0 indicates otherwise. We perform speaker identification and verification using a missing data recognizer under cochannel and other noise conditions, and show that the ideal binary mask provides large performance gains. By employing a speech segregation system that estimates the ideal binary mask, we achieve significant improvements over alternative approaches. Our study, thus, demonstrates that the use of binary masking represents a promising direction for robust speaker recognition. DeLiang Wang |
ICASSP (1) | 2 |
| 2006 | A Supervised Learning Approach to Uncertainty Decoding for Robust Speech RecognitionabstractRecently several algorithms have been proposed to enhance noisy speech by estimating a binary mask that can be used to select those time-frequency regions of a noisy speech signal that contain more speech energy than noise energy. This binary mask encodes the uncertainty associated with enhanced speech in the linear spectral domain. The use of the cepstral transformation leads to a smearing of this uncertainty. We propose a supervised approach to learn the non linear transformation of the uncertainty from the linear spectral domain to the cepstral domain. This uncertainty is used by a decoder that exploits the variance associated with the enhanced cepstral features to improve robust speech recognition. Systematic evaluations on a subset of the Aurora4 task using the estimated uncertainty shows substantial improvement over the baseline performance. Soundararajan Srinivasan, DeLiang Wang |
ICASSP (1) | 2 |
| 2006 | Unvoiced Speech Segregationabstractspeech segregation, or the cocktail party problem, has proven to be extremely challenging. While efforts in computational auditory scene analysis have led to considerable progress in voiced speech which lacks harmonic structure and has weaker energy, hence more susceptible to interference. We describe a novel approach to address this problem. The segregation process occurs in two stages: segmentation and grouping. In segmentation, our model decomposes the input mixture into contiguous time-frequency segments by analyzing sound onsets and offsets. Grouping of unvoiced segments is based on Bayesian classification of acousticphonetic features. The proposed model yields very promising results. DeLiang Wang, Guoning Hu |
ICASSP (5) | 1 |
| 2006 | A computational auditory scene analysis system for robust speech recognitionabstractWe present a computational auditory scene analysis system for separating and recognizing target speech in the presence of competing speech or noise. We estimate, in two stages, the ideal binary time-frequency (T-F) mask which retains the mixture in a local T-F unit if and only if the target is stronger than the interference within the unit. In the first stage, we use harmonicity to segregate the voiced portions of individual sources in each time frame based on multipitch tracking. Additionally, unvoiced portions are segmented based on an onset/offset analysis. In the second stage, speaker characteristics are used to group the T-F units across time frames. The resulting T-F masks are used in conjunction with missing-data methods for recognition. Systematic evaluations on a speech separation challenge task show significant improvement over the baseline performance. Index Terms: speech segregation, computational auditory scene analysis, binary time-frequency mask, robust speech recognition. Soundararajan Srinivasan, Zhaozhang Jin, DeLiang Wang |
INTERSPEECH | 4 |
| 2006 | Binary and ratio time-frequency masks for robust speech recognition
Soundararajan Srinivasan, Nicoleta Roman, DeLiang Wang |
Speech Commun. | 3 |
| 2006 | Model-based sequential organization in cochannel speechabstractA human listener has the ability to follow a speaker's voice while others are speaking simultaneously; in particular, the listener can organize the time-frequency energy of the same speaker across time into a single stream. In this paper, we focus on sequential organization in cochannel speech, or mixtures of two voices. We extract minimally corrupted segments, or usable speech, in cochannel speech using a robust multipitch tracking algorithm. The extracted usable speech is shown to capture speaker characteristics and improves speaker identification (SID) performance across various target-to-interferer ratios. To utilize speaker characteristics for sequential organization, we extend the traditional SID framework to cochannel speech and derive a joint objective for sequential grouping and SID, leading to a problem of search for the optimum hypothesis. Subsequently we propose a hypothesis pruning algorithm based on speaker models in order to make the search computationally efficient. Evaluation results show that the proposed system approaches the ceiling SID performance obtained with prior pitch information and yields significant improvement over alternative approaches to sequential organization. DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | A two-stage algorithm for one-microphone reverberant speech enhancementabstractUnder noise-free conditions, the quality of reverberant speech is dependent on two distinct perceptual components: coloration and long-term reverberation. They correspond to two physical variables: signal-to-reverberant energy ratio (SRR) and reverberation time, respectively. Inspired by this observation, we propose a two-stage reverberant speech enhancement algorithm using one microphone. In the first stage, an inverse filter is estimated to reduce coloration effects or increase SRR. The second stage employs spectral subtraction to minimize the influence of long-term reverberation. The proposed algorithm significantly improves the quality of reverberant speech. A comparison with a recent enhancement algorithm is made on a corpus of speech utterances in a number of reverberant conditions, and the results show that our algorithm performs substantially better. DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Separation of Fricatives and AffricatesabstractSeparating speech from acoustic interference is a very challenging task. In particular, no system successfully addresses the separation of unvoiced speech. Fricatives and affricates are two main categories of consonants that contain a significant amount of unvoiced signal. We propose a novel system that separates fricatives and affricates from non-speech interference. The system first decomposes the input mixture into segments, each of which contains signal mainly from one source. Then it detects segments dominated by unvoiced portions of fricatives and affricates with a feature-based Bayesian classifier, and groups these segments with voiced speech separated by a previous system. The proposed system is evaluated with various types of interference and produces promising results. Guoning Hu, DeLiang Wang |
ICASSP (1) | 2 |
| 2005 | Detecting pitch of singing voice in polyphonic audioabstractWe propose a robust algorithm to detect the pitch of a singing voice in polyphonic audio. A new channel/peak selection scheme is introduced to exploit the salience of the singing voice and the beating phenomenon in high frequency channels. An HMM is employed to integrate the periodicity information across frequency channels and time frames. Quantitative evaluation shows that the new system performs significantly better than existing algorithms for predominant pitch detection in polyphonic audio. DeLiang Wang |
ICASSP (3) | 2 |
| 2005 | Robust Speech Recognition by Integrating Speech Separation and Hypothesis TestingabstractMissing data methods attempt to improve robust speech recognition by distinguishing between reliable and unreliable data in the time-frequency domain. Such methods require a binary mask which labels time-frequency regions of a noisy speech signal as reliable if they contain more speech energy than noise energy and unreliable otherwise. Current methods for estimating the mask are based mainly on bottom-up speech separation cues such as harmonicity and produce labeling errors that cause a degradation in recognition performance. We propose a two stage recognition system in order to improve mask estimation and produce better recognition results. First, an n-best lattice consistent with the speech separation mask is generated. The lattice is then re-scored by expanding the mask using a model-based hypothesis test to determine the reliability of individual time-frequency regions. Systematic evaluations show significant improvement in recognition performance compared to that using speech separation. Soundararajan Srinivasan, DeLiang Wang |
ICASSP (1) | 2 |
| 2005 | A Two-Stage Algorithm for Enhancement of Reverberant SpeechabstractRoom reverberation causes two perceptual distortions on clean speech: coloration and long-term reverberation. These two effects correspond to two physical variables: signal-to-reverberant energy ratio (SRR) and reverberation time, respectively. Based on this observation, we propose a two-stage algorithm that enhances reverberant speech from one-microphone recordings. In the first stage, an inverse filter is estimated to reduce coloration effects or increase SRR. The second stage employs spectral subtraction to minimize the influence of long-term reverberation. The proposed algorithm significantly improves the quality of reverberant speech. A comparison with a recent one-microphone enhancement algorithm shows that our system produces significantly better results. DeLiang Wang |
ICASSP (1) | 2 |
| 2005 | A pitch-based model for separation of reverberant speechabstractIn everyday listening, both background noise and reverberation degrade the speech signal. While monaural speech separation based on periodicity has achieved considerable progress in handling additive noise, little research has been devoted to reverberant scenarios. Reverberation smears the harmonic structure of speech signals, and our evaluations using a pitch-based separation algorithm show that an increase in the room reverberation time causes degradation in performance due to the loss in periodicity for the target signal. We propose a two-stage monaural speech separation system that combines the inverse filtering of the room impulse response corresponding to target location with a pitch-based speech segregation method. As a result of the first processing stage, the harmonicity of a signal arriving from target direction is partially restored while signals arriving from other locations are further smeared, and this leads to improved separation. A systematic evaluation shows that the proposed system results in considerable signal-to-noise ratio gains across different conditions. 1. Nicoleta Roman, DeLiang Wang |
INTERSPEECH | 2 |
| 2005 | Modeling the perception of multitalker speechabstractListeners ’ ability to understand a target speaker in the presence of one or more simultaneous competing speakers is subject to two types of masking: Energetic and informational. Energetic masking occurs when target and interfering signals overlap in time and frequency resulting in portions of target becoming inaudible. Informational masking occurs when the listener is unable to segregate the target from interference, while both are audible. We present a model of multitalker speech perception that accounts for both types of masking. Human perception in the presence of energetic masking is modeled using a speech recognizer that treats the masked time-frequency units of target as missing data. The effects of informational masking on the recognizer are modeled using the output of a speech segregation system. On a systematic evaluation, the performance of the proposed model is in broad agreement with perceptual results. 1. Soundararajan Srinivasan, DeLiang Wang |
INTERSPEECH | 2 |
| 2005 | A schema-based model for phonemic restoration
Soundararajan Srinivasan, DeLiang Wang |
Speech Commun. | 2 |
| 2005 | The time dimension for scene analysisabstractA fundamental issue in neural computation is the binding problem, which refers to how sensory elements in a scene organize into perceived objects, or percepts. The issue of binding is hotly debated in recent years in neuroscience and related communities. Much of the debate, however, gives little attention to computational considerations. This review intends to elucidate the computational issues that bear directly on the binding issue. The review starts with two problems considered by Rosenblatt to be the most challenging to the development of perceptron theory more than 40 years ago, and argues that the main challenge is the figure-ground separation problem, which is intrinsically related to the binding problem. The theme of the review is that the time dimension is essential for systematically attacking Rosenblatt's challenge. The temporal correlation theory as well as its special form--oscillatory correlation theory-is discussed as an adequate representation theory to address the binding problem. Recent advances in understanding oscillatory dynamics are reviewed, and these advances have overcome key computational obstacles for the development of the oscillatory correlation theory. We survey a variety of studies that address the scene analysis problem. The results of these studies have substantially advanced the capability of neural networks for figure-ground separation. A number of issues regarding oscillatory correlation are considered and clarified. Finally, the time dimension is argued to be necessary for versatile computing. DeLiang Wang |
IEEE Trans. Neural Networks | 1 |
| 2004 | Binaural sound segregation for multisource reverberant environmentsabstractWe present a novel method for binaural sound segregation from acoustic mixtures contaminated by both multiple interference and reverberation. We employ the notion of an ideal time-frequency binary mask, which selects the target if it is stronger than the interference in a local time-frequency (T-F) unit. As opposed to classical adaptive filtering, which focuses on the suppression of noise, our model employs an adaptive filter that performs target cancellation. T-F units dominated by a target are largely suppressed at the output of the cancellation unit when compared to units dominated by noise. Consequently, the actual input-to-output attenuation level in each T-F unit is used to estimate an ideal binary mask. A systematic evaluation in terms of automatic speech recognition performance shows that the resulting system produces masks close to ideal binary ones. Nicoleta Roman, DeLiang Wang |
ICASSP (2) | 2 |
| 2004 | A comparison of CNN and LEGION networksabstractCNN and LEGION networks have been extensively studied in recent years. These two frameworks share many common features; both employ continuous-time dynamics, are nonlinear, and emphasize local connectivity. In addition, they both have been successfully applied to visual processing tasks and implemented on analog VLSI chips. This paper investigates the relations between the two frameworks. We present their standard versions, and contrast the underlying dynamics and connectivity. We also describe several tasks where both CNN and LEGION have been applied. The comparison reveals fundamental differences between them. CNN is good for early visual processing, whereas LEGION is good for midlevel visual processing. Furthermore, the comparison suggests that a combined network is likely to enhance the overall processing capability. DeLiang Wang |
IJCNN | 1 |
| 2004 | Model-based sequential organization for cochannel speaker identification
DeLiang Wang |
INTERSPEECH | 2 |
| 2004 | On binary and ratio time-frequency masks for robust speech recognition
Soundararajan Srinivasan, Nicoleta Roman, DeLiang Wang |
INTERSPEECH | 3 |
| 2004 | A binaural processor for missing data speech recognition in the presence of noise and small-room reverberation
Kalle J. Palomäki, Guy J. Brown, DeLiang Wang |
Speech Commun. | 3 |
| 2004 | Synchronization rates in classes of relaxation oscillatorsabstractRelaxation oscillators arise frequently in physics, electronics, mathematics, and biology. Their mathematical definitions possess a high degree of flexibility in the sense that through appropriate parameter choices relaxation oscillators can be made to exhibit qualitatively different kinds of oscillations. We study numerically four different classes of relaxation oscillators through their synchronization rates in one-dimensional chains with a Heaviside step function interaction and obtain the following results. Relaxation oscillators in the sinusoidal and relaxation regime both exhibit an average time to synchrony, approximately n, where n is the chain length. Relaxation oscillators in the singular limit exhibit approximately n(p), where p is a numerically obtained value less than 0.5. Relaxation oscillators in the singular limit with parameters modified so that they resemble spike oscillations exhibit approximately log(n) in chains and approximately log(L) in two-dimensional square networks of length L. Finally, using a sigmoid interaction results in approximately n(2), for relaxation oscillators in the sinusoidal and relaxation regimes, indicating that the form of the coupling is a controlling factor in the synchronization rate. Shannon R. Campbell, DeLiang Wang, Ciriyam Jayaprakash |
IEEE Trans. Neural Networks | 2 |
| 2004 | Monaural speech segregation based on pitch tracking and amplitude modulationabstractSegregating speech from one monaural recording has proven to be very challenging. Monaural segregation of voiced speech has been studied in previous systems that incorporate auditory scene analysis principles. A major problem for these systems is their inability to deal with the high-frequency part of speech. Psychoacoustic evidence suggests that different perceptual mechanisms are involved in handling resolved and unresolved harmonics. We propose a novel system for voiced speech segregation that segregates resolved and unresolved harmonics differently. For resolved harmonics, the system generates segments based on temporal continuity and cross-channel correlation, and groups them according to their periodicities. For unresolved harmonics, it generates segments based on common amplitude modulation (AM) in addition to temporal continuity and groups them according to AM rates. Underlying the segregation process is a pitch contour that is first estimated from speech segregated according to dominant pitch and then adjusted according to psychoacoustic constraints. Our system is systematically evaluated and compared with pervious systems, and it yields substantially better performance, especially for the high-frequency part of speech. Guoning Hu, DeLiang Wang |
IEEE Trans. Neural Networks | 2 |
| 2003 | Separation of stop consonantsabstractTo extract speech from acoustic interference is a challenging problem. Previous systems based on auditory scene analysis principles deal with voiced speech, but cannot separate unvoiced speech. We propose a novel method to separate stop consonants, which contain significant unvoiced signals, based on their acoustic properties. The method employs onset as the major grouping cue; it first detects stops through onset detection and feature-based Bayesian classification, then groups detected onsets based on onset coincidence. This method is tested with utterances mixed with various types of interference. Guoning Hu, DeLiang Wang |
ICASSP (2) | 2 |
| 2003 | Binaural tracking of multiple moving sourcesabstractThis paper presents a novel method for tracking the azimuth locations of multiple active sources based on binaural processing. Binaural cues are strongly correlated with source locations for spectral regions dominated by only one source. Therefore, this approach integrates reliable information across different frequency channels to produce a likelihood function in the target space. Finally, a hidden Markov model (HMM) is employed for forming continuous tracks and detecting the number of active sources across time. Experimental results are presented for simulated multi-source scenarios. Nicoleta Roman, DeLiang Wang |
ICASSP (5) | 2 |
| 2003 | Co-channel speaker identification using usable speech extraction based on multi-pitch trackingabstractRecently, usable speech criteria have been proposed to extract minimally corrupted speech for speaker identification (SID) in co-channel speech. In this paper, we propose a new usable speech extraction method to improve the SID performance under the co-channel situation based on the pitch information obtained from a robust multi-pitch tracking algorithm [2]. The idea is to retain the speech segments that have only one pitch detected and remove the others. The system is evaluated on co-channel speech and results show a significant improvement across various target to interferer ratios (TIR) for speaker identification. DeLiang Wang |
ICASSP (2) | 2 |
| 2003 | A one-microphone algorithm for reverberant speech enhancementabstractWe present an algorithm for reverberant speech enhancement using one microphone. We first propose a novel pitch-based reverberation measure for estimating reverberation time (RT60) based on the distribution of relative time lags. This measure of pitch strength correlates with reverberation and decreases systematically as detrimental effects of reverberation on harmonic structure increase. Then a reverberant speech enhancement method is developed to estimate and subtract later echo components. The results show that our approach appreciably reduces reverberation effects. DeLiang Wang |
ICASSP (1) | 2 |
| 2003 | On intrinsic generalization of low dimensional representations of images for recognitionabstractLow dimensional representations of images impose equivalence relations in the image space; the induced equivalence class of an image is named as its intrinsic generalization. The intrinsic generalization of a representation provides a novel way to measure its generalization and leads to more fundamental insights than the commonly used recognition performance, which is heavily influenced by the choice of training and test data. We demonstrate the limitations of linear subspace representations by sampling their intrinsic generalization, and propose a nonlinear representation that overcomes these limitations. The proposed representation projects images nonlinearly into the marginal densities of their filter responses, followed by linear projections of the marginals. We have used experiments on large datasets to show that the representations that have better intrinsic generalization also lead to a better recognition performance. Xiuwen Liu 0001, Anuj Srivastava, DeLiang Wang |
IJCNN | 3 |
| 2003 | Monaural speech segregation and oscillatory correlationabstractSummary form only given. Speech segregation from a monaural recording is a primary task of auditory grouping, and has proven to be very challenging. Theoretical and empirical investigations of brain functions point to the mechanism of oscillatory correlation as a plausible framework for perceptual grouping. In this framework, an assembly of synchronized oscillators represents a stream, and oscillator assemblies that desynchronize from one another represent different groups. We describe a multi-stage model for the monaural speech segregation task. The model starts with simulated auditory periphery. A subsequent stage computes mid-level auditory representations, including correlograms and cross-channel correlations. Underlying auditory segmentation and grouping is a neural oscillator network that implements oscillatory correlation. The network encodes proximity in frequency and time, periodicity, and amplitude modulation (AM). Motivated by psychoacoustic observations, our system employs different mechanism to handle resolved and unresolved harmonics. The model has been systematically evaluated, and it yields substantially better performance than previous systems. DeLiang Wang |
IJCNN | 1 |
| 2003 | Schema-based modeling of phonemic restoration
Soundararajan Srinivasan, DeLiang Wang |
INTERSPEECH | 2 |
| 2003 | A Classification-based Cocktail-party ProcessorabstractGuy J. Brown Department of Computer Science University of Sheffield 211 Portobello Street Sheffield, S1 4DP, UK [email protected] Nicoleta Roman, DeLiang Wang, Guy J. Brown |
NIPS | 2 |
| 2003 | Intrinsic generalization analysis of low dimensional representations
Xiuwen Liu 0001, Anuj Srivastava, DeLiang Wang |
Neural Networks | 3 |
| 2003 | Welcome to the special issue: the best of the best
Donald C. Wunsch II, Michael E. Hasselmo, DeLiang Wang, Ganesh K. Venayagamoorthy |
Neural Networks | 3 |
| 2003 | A multipitch tracking algorithm for noisy speechabstractAn effective multipitch tracking algorithm for noisy speech is critical for acoustic signal processing. However, the performance of existing algorithms is not satisfactory. We present a robust algorithm for multipitch tracking of noisy speech. Our approach integrates an improved channel and peak selection method, a new method for extracting periodicity information across different channels, and a hidden Markov model (HMM) for forming continuous pitch tracks. The resulting algorithm can reliably track single and double pitch tracks in a noisy environment. We suggest a pitch error measure for the multipitch situation. The proposed algorithm is evaluated on a database of speech utterances mixed with various types of interference. Quantitative comparisons show that our algorithm significantly outperforms existing ones. DeLiang Wang, Guy J. Brown |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Texture classification using spectral histogramsabstractBased on a local spatial/frequency representation,we employ a spectral histogram as a feature statistic for texture classification. The spectral histogram consists of marginal distributions of responses of a bank of filters and encodes implicitly the local structure of images through the filtering stage and the global appearance through the histogram stage. The distance between two spectral histograms is measured using chi(2)-statistic. The spectral histogram with the associated distance measure exhibits several properties that are necessary for texture classification. A filter selection algorithm is proposed to maximize classification performance of a given dataset. Our classification experiments using natural texture images reveal that the spectral histogram representation provides a robust feature statistic for textures and generalizes well. Comparisons show that our method produces a marked improvement in classification performance. Finally we point out the relationships between existing texture features and the spectral histogram, suggesting that the latter may provide a unified texture feature. Xiuwen Liu 0001, DeLiang Wang |
IEEE Trans. Image Process. | 2 |
| 2002 | Monaural speech segregation based on pitch tracking and amplitude modulationabstractMonaural speech segregation remains a computational challenge for auditory scene analysis (ASA). A major problem for existing computational auditory scene analysis (CASA) systems is their inability to deal with signals in the high-frequency range. Psychoacoustic evidence suggests that different perceptual mechanisms are involved to handle resolved and unresolved harmonics. We propose a system for speech segregation that deals with low-frequency and high-frequency signals differently. For low-frequency signals, our model generates segments based on temporal continuity and cross-channel correlation, and groups them according to periodicity. For high-frequency signals. the model generates segments based on common amplitude modulation (AM) in addition to temporal continuity, and groups them according to AM repetition rates. Underlying the grouping process is a pitch contour that is first estimated from segregated speech based on global pitch and then verified by psychoacoustic constraints. Our system is systematically evaluated, and it yields substantially better performance than previous CASA systems, especially in the high-frequency range. Guoning Hu, DeLiang Wang |
ICASSP | 2 |
| 2002 | Location-based sound segregationabstractAt a cocktail party, we can selectively attend to a single voice and filter out all the other acoustical interferences. How to simulate this perceptual ability remains a great challenge. This paper describes a novel location-based approach for speech segregation. The auditory masking effect motivates the notion of an “ideal” time-frequency binary mask, which selects the target if it is stronger than the interference in a local time-frequency region. We observe that within a narrow frequency band modifications to the relative energy of the target source with respect to the interfering energy trigger systematic deviations for binaural cues. For a given spatial configuration, this interaction produces characteristic clustering in the binaural feature space. Consequently, we perform pattern classification in order to estimate ideal binary masks. A systematic evaluation shows that the resulting system produces masks very close to ideal binary ones, and large improvement over previous models. Nicoleta Roman, DeLiang Wang, Guy J. Brown |
ICASSP | 2 |
| 2002 | A multi-pitch tracking algorithm for noisy speechabstractWe present a robust algorithm for multi-pitch tracking of noisy speech. Our approach integrates an improved channel and peak selection method, a new integration method for extracting periodicity information across different frequency channels, and a hidden Markov model (HMM) for forming continuous pitch tracks, and as a result, our algorithm can reliably track single and double pitch tracks in a noisy environment. The proposed algorithm is evaluated on a database of speech utterances mixed with various interferences and the results show that our algorithm outperforms existing algorithms significantly. DeLiang Wang, Guy J. Brown |
ICASSP | 2 |
| 2002 | Monaural Speech Separationabstractthat deals with Monaural speech separation has been studied in previous systems that incorporate auditory scene analysis principles. A major problem for these systems is their inability to deal with speech in the high- frequency range. Psychoacoustic evidence suggests that different perceptual mechanisms are involved in handling resolved and unresolved harmonics. Motivated by this, we propose a model for monaural separation low-frequency and high- frequency signals differently. For resolved harmonics, our model generates segments based on temporal continuity and cross-channel correlation, and groups them according to periodicity. For unresolved harmonics, the model generates segments based on amplitude modulation (AM) in addition to temporal continuity and groups them according to AM repetition rates derived from sinusoidal modeling. Underlying the separation process is a pitch contour obtained according to psychoacoustic constraints. Our model is systematically evaluated, and it yields substantially better performance than previous systems, especially in the high-frequency range. Guoning Hu, DeLiang Wang |
NIPS | 2 |
| 2002 | A dynamically coupled neural oscillator network for image segmentation
Ke Chen 0001, DeLiang Wang |
Neural Networks | 2 |
| 2002 | Scene analysis by integrating primitive segmentation and associative memoryabstractScene analysis is a major aspect of perception and continues to challenge machine perception. This paper addresses the scene-analysis problem by integrating a primitive segmentation stage with a model of associative memory. The model is a multistage system that consists of an initial primitive segmentation stage, a multimodule associative memory, and a short-term memory (STM) layer. Primitive segmentation is performed by a locally excitatory globally inhibitory oscillator network (LEGION), which segments the input scene into multiple parts that correspond to groups of synchronous oscillations. Each segment triggers memory recall and multiple recalled patterns then interact with one another in the STM layer. The STM layer projects to the LEGION network, giving rise to memory-based grouping and segmentation. The system achieves scene analysis entirely in phase space, which provides a unifying mechanism for both bottom-up analysis and top-down analysis. The model is evaluated with a systematic set of three-dimensional (3-D) line drawing objects, which are arranged in an arbitrary fashion to compose input scenes that allow object occlusion. Memory-based organization is responsible for a significant improvement in performance. A number of issues are discussed, including input-anchored alignment, top-down organization, and the role of STM in producing context sensitivity of memory recall. DeLiang Wang, Xiuwen Liu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2001 | Image segmentation using local spectral histogramsabstractWe propose a new algorithm for image segmentation. We use the spectral histogram, which is a vector consisting of marginal distributions of responses from chosen filters as a generic feature for texture as well as intensity images. Motivated by a new segmentation energy functional, we derive an iterative and deterministic approximation algorithm for segmentation. Based on the relationships between different scales and neighboring windows, we also develop an algorithm which can automatically detect homogeneous regions in an input image, which may consist of texture regions. To reduce the boundary uncertainty due to the large spatial window used for spectral histograms, we propose a novel local feature by building precise probability models based on current segmentation results. We have applied our algorithm to intensity, texture, and natural images and obtained good results with accurate texture boundaries. Xiuwen Liu 0001, DeLiang Wang, Anuj Srivastava |
ICIP (1) | 2 |
| 2001 | Synchronization in Relaxation Oscillator Networks with Conduction DelaysabstractWe study locally coupled networks of relaxation oscillators with excitatory connections and conduction delays and propose a mechanism for achieving zero phase-lag synchrony. Our mechanism is based on the observation that different rates of motion along different nullclines of the system can lead to synchrony in the presence of conduction delays. We analyze the system of two coupled oscillators and derive phase compression rates. This analysis indicates how to choose nullclines for individual relaxation oscillators in order to induce rapid synchrony. The numerical simulations demonstrate that our analytical results extend to locally coupled networks with conduction delays and that these networks can attain rapid synchrony with appropriately chosen nullclines and initial conditions. The robustness of the proposed mechanism is verified with respect to different nullclines, variations in parameter values, and initial conditions. Jeffrey J. Fox, Ciriyam Jayaprakash, DeLiang Wang, Shannon R. Campbell |
Neural Comput. | 3 |
| 2001 | A comparison of auditory and blind separation techniques for speech segregationabstractA fundamental problem in auditory and speech processing is the segregation of speech from concurrent sounds. This problem has been a focus of study in computational auditory scene analysis (CASA), and it has also been investigated from the perspective of blind source separation. Using a standard corpus of voiced speech mixed with interfering sounds, we report a comparison between CASA and blind source separation techniques, which have been developed independently. Our comparison reveals that they perform well under very different conditions. A number of conclusions are drawn with respect to their relative strengths and weaknesses in speech segregation applications as well as in modeling auditory function. André J. W. van der Kouwe, DeLiang Wang, Guy J. Brown |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Extraction of hydrographic regions from remote sensing images using an oscillator network with weight adaptationabstractThe authors propose a framework for object extraction with accurate boundaries. A multilayer perceptron is used to identify seed points through examples, and regions are extracted and localized using a locally coupled network with weight adaptation. A functional system has been developed and applied to hydrographic region extraction from Digital Orthophoto Quarter-Quadrangle images. Xiuwen Liu 0001, Ke Chen 0001, DeLiang Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2001 | Texture segmentation using Gaussian-Markov random fields and neural oscillator networksabstractWe propose an image segmentation method based on texture analysis. Our method is composed of two parts. The first part determines a novel set of texture features derived from a Gaussian-Markov random fields (GMRF) model. Unlike a GMRF-based approach, our method does not employ model parameters as features or require the extraction of features for a fixed set of texture types a priori. The second part is a 2D array of locally excitatory globally inhibitory oscillator networks (LEGION). After being filtered for noise suppression, features are used to determine the local couplings in the network. When LEGION runs, the oscillators corresponding to the same texture tend to synchronize, whereas different texture regions tend to correspond to distinct phases. In simulations, a large system of differential equations is solved for the first time using a recently proposed method for integrating relaxation oscillator networks. We provide results on real texture images to demonstrate the performance of our method. Erdogan Çesmeli, DeLiang Wang |
IEEE Trans. Neural Networks | 2 |
| 2001 | Perceiving geometric patterns: from spirals to inside-outside relationsabstractSince first proposed by Minsky and Papert (1969), the spiral problem is well known in neural networks. It receives much attention as a benchmark for various learning algorithms. Unlike previous work that emphasizes learning, we approach the problem from a different perspective. We point out that the spiral problem is intrinsically connected to the inside-outside problem proposed by Ullman (1984, 1996). We propose a solution to both problems based on oscillatory correlation using a time-delay network. Our simulation results are qualitatively consistent with human performance, and we interpret human limitations in terms of synchrony and time delays. As a special case, our network without time delays can always distinguish these figures regardless of shape, position, size, and orientation. Ke Chen 0001, DeLiang Wang |
IEEE Trans. Neural Networks | 2 |
| 2000 | An Oscillatory Correlation Model of Human Motion PerceptionabstractAn oscillatory correlation model of human motion perception is proposed based on the integration of motion and luminance information. The model is composed of two parallel pathways that segment the input scene based on motion and luminance, respectively. Combining these segmentations, the model refines the motion estimates in the integration stage to obtain the final segmentation in the motion pathway. For segmentation, LEGION (locally excitatory globally inhibitory oscillator network) is employed whereby the phases of oscillators are used for region labeling. The model performance is demonstrated using a set of psychophysical data. Erdogan Çesmeli, Delwin T. Lindsey, DeLiang Wang |
IJCNN (4) | 3 |
| 2000 | On Connectedness: A Solution Based on Oscillatory CorrelationabstractA long-standing problem in Neural Comp has been the problem of connectedness, first identified by Minsky and Papert (1969). This problem served as the cornerstone for them to establish analytically that perceptrons are fundamentally limited in computing geometrical (topological) properties. A solution to this problem is offered by a different class of neural networks: oscillator networks. To solve the problem, the representation of oscillatory correlation is employed, whereby one pattern is represented as a synchronized block of oscillators and different patterns are represented by distinct blocks that desynchronize from each other. Oscillatory correlation emerges from LEGION (locally excitatory globally inhibitory oscillator network), whose architecture consists of local excitation and global inhibition among neural oscillators. It is further shown that these oscillator networks exhibit sensitivity to topological structure, which may lay a neurocomputational foundation for explaining the psychophysical phenomenon of topological perception. DeLiang Wang |
Neural Comput. | 1 |
| 2000 | Boundary detection by contextual non-linear smoothing
Xiuwen Liu 0001, DeLiang Wang, J. Raul Ramirez |
Pattern Recognit. | 2 |
| 2000 | Motion segmentation based on motion/brightness integration and oscillatory correlationabstractA segmentation method based on the integration of motion and brightness is proposed for image sequences. The method is composed of two parallel pathways that process motion and brightness, respectively. Inspired by the visual system, the motion pathway has two stages. The first stage estimates local motion at locations with reliable information. The second stage performs segmentation based on local motion estimates. In the brightness pathway, the input scene is segmented into regions based on brightness distribution. Subsequently, segmentation results from the two pathways are integrated to refine motion estimates. The final segmentation is performed in the motion network based on refined estimates. For segmentation, locally excitatory globally inhibitory oscillator network (LEGION) architecture is employed whereby the oscillators corresponding to a region of similar motion/brightness oscillate in synchrony and different regions attain different phases. Results on synthetic and real image sequences are provided, and comparisons with other methods are made. Erdogan Çesmeli, DeLiang Wang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2000 | Weight adaptation and oscillatory correlation for image segmentationabstractWe propose a method for image segmentation based on a neural oscillator network. Unlike previous methods, weight adaptation is adopted during segmentation to remove noise and preserve significant discontinuities in an image. Moreover, a logarithmic grouping rule is proposed to facilitate grouping of oscillators representing pixels with coherent properties. We show that weight adaptation plays the roles of noise removal and feature preservation. In particular, our weight adaptation scheme is insensitive to termination time and the resulting dynamic weights in a wide range of iterations lead to the same segmentation results. A computer algorithm derived from oscillatory dynamics is applied to synthetic and real images and simulation results show that the algorithm yields favorable segmentation results in comparison with other recent algorithms. In addition, the weight adaptation scheme can be directly transformed to a novel feature-preserving smoothing procedure. We also demonstrate that our nonlinear smoothing algorithm achieves good results for various kinds of images. Ke Chen 0001, DeLiang Wang, Xiuwen Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 1999 | Image segmentation based on a dynamically coupled neural oscillator networkabstractIn this paper, a dynamically coupled neural oscillator network is proposed for image segmentation. Instead of pair-wise coupling, an ensemble of oscillators coupled in a local region is used for grouping. We introduce a set of neighborhoods to generate dynamical coupling structures associated with a specific oscillator. Based on the proximity and similarity principles, two grouping rules are proposed to explicitly consider the distinct cases of whether an oscillator is inside a homogeneous image region or near a boundary between different regions. The use of dynamical coupling makes our segmentation network robust to noise on an image. For fast computation, a segmentation algorithm is abstracted from the underlying oscillatory dynamics and has been applied to synthetic and real images. Simulation results demonstrate the effectiveness of our oscillator network in image segmentation. Ke Chen 0001, DeLiang Wang |
IJCNN | 2 |
| 1999 | The separation of speech from interfering sounds: an oscillatory correlation approachabstractA neural model is described which uses oscillatory correlation to segregate speech from interfering sound sources. The core of the model is a two-layer neural oscillator network. The first layer of the network identifies the connected regions of energy in the time-frequency plane (segments). In the second layer, segments that have a common fundamental frequency are grouped into streams. A stream is represented by a synchronized population of relaxation oscillators, and different streams are represented by desynchronized oscillator populations. The model has been evaluated using a corpus of voiced speech mixed with interfering sounds, and produces an improvement in signal-to-noise ratio for every mixture. Guy J. Brown, DeLiang Wang |
IJCNN | 2 |
| 1999 | Image segmentation based on motion/luminance integration and oscillatory correlationabstractAn image segmentation method is proposed based on the integration of motion and luminance information. The method is composed of two parallel pathways that process motion and luminance, respectively. Inspired by the visual system, the motion pathway has two stages. The first stage estimates local motion at locations with reliable information The second stage groups locations based on their motion estimates. In the parallel pathway, the input scene is segmented based on luminance. In the subsequent integration stage, motion estimates are refined to obtain the final segmentation result in the motion pathway. For segmentation, LEGION (Locally Excitatory Globally Inhibitory Oscillator Networks) is employed whereby the phases of oscillators are used for region labeling. Results on synthetic and real image sequences are provided. Erdogan Çesmeli, DeLiang Wang |
IJCNN | 2 |
| 1999 | A boundary-pair representation for perception modelingabstractIt is widely accepted that responses from on- and off-center cells give rise to edges and are equivalent to edge detectors. In this paper, we point out that on- and off-center cell responses provide more information than edges. We show that an edge-based representation makes the ownership of boundaries ambiguous and requires a combinatorial search to model perceptual grouping. By analyzing the differences between edges and responses from on- and off-center cells, we propose a boundary-pair representation, which makes the ownership of boundaries explicit and eliminates the need of a combinatorial search computationally. Each boundary in the boundary-pair representation is associated with regional attributes. We show that this representation is equivalent to a surface representation through a local diffusion. This provides a unified representation for perception modeling. Based on this representation, a figure-ground segregation network is constructed to demonstrate the capabilities of the model in explaining many perceptual phenomena. Xiuwen Liu 0001, DeLiang Wang |
IJCNN | 2 |
| 1999 | Perceptual organization based on temporal dynamicsabstractThis paper presents a computational model for perceptual organization. A figure-ground segregation network is proposed based on a novel boundary pair representation. The system solves the figure-ground segregation problem through temporal evolution. Gestalt-like grouping rules are incorporated by modulating connections, which determines the temporal behavior and thus the perception of the system. The results are then fed to a surface completion module based on local diffusion. Different perceptual phenomena, such as modal and a modal completion, virtual contours, grouping and shape decomposition are explained by the model with a fixed set of parameters. Computationally, the system eliminates combinatorial optimization, which is common to many existing computational approaches. It also accounts for more examples that are consistent with psychological experiments. In addition, the boundary-pair representation is consistent with well-known on- and off-center cell responses and thus biologically more plausible. Xiuwen Liu 0001, DeLiang Wang |
IJCNN | 2 |
| 1999 | An Oscillatory Correlation Frame work for Computational Auditory Scene Analysis
Guy J. Brown, DeLiang Wang |
NIPS | 2 |
| 1999 | Perceptual Organization Based on Temporal Dynamics
Xiuwen Liu 0001, DeLiang Wang |
NIPS | 2 |
| 1999 | Synchrony and Desynchrony in Integrate-and-Fire OscillatorsabstractDue to many experimental reports of synchronous neural activity in the brain, there is much interest in understanding synchronization in networks of neural oscillators and its potential for computing perceptual organization. Contrary to Hopfield and Herz (1995), we find that networks of locally coupled integrate-and-fire oscillators can quickly synchronize. Furthermore, we examine the time needed to synchronize such networks. We observe that these networks synchronize at times proportional to the logarithm of their size, and we give the parameters used to control the rate of synchronization. Inspired by locally excitatory globally inhibitory oscillator network (LEGION) dynamics with relaxation oscillators (Terman & Wang, 1995), we find that global inhibition can play a similar role of desynchronization in a network of integrate-and-fire oscillators. We illustrate that a LEGION architecture with integrate-and-fire oscillators can be similarly used to address image analysis. Shannon R. Campbell, DeLiang Wang, Ciriyam Jayaprakash |
Neural Comput. | 2 |
| 1999 | Object selection based on oscillatory correlation
DeLiang Wang |
Neural Networks | 1 |
| 1999 | Segmentation of Medical Images Using LEGIONabstractAdvances in visualization technology and specialized graphic workstations allow clinicians to virtually interact with anatomical structures contained within sampled medical-image datasets. A hindrance to the effective use of this technology is the difficult problem of image segmentation. In this paper, we utilize a recently proposed oscillator network called the locally excitatory globally inhibitory oscillator network (LEGION) whose ability to achieve fast synchrony with local excitation and desynchrony with global inhibition makes it an effective computational framework for grouping similar features and segregating dissimilar ones in an image. We extract an algorithm from LEGION dynamics and propose an adaptive scheme for grouping. We show results of the algorithm to two-dimensional (2-D) and three-dimensional (3-D) (volume) computerized topography (CT) and magnetic resonance imaging (MRI) medical-image datasets. In addition, we compare our algorithm with other algorithms for medical-image segmentation, as well as with manual segmentation. LEGION's computational and architectural properties make it a promising approach for real-time medical-image segmentation. Naeem Shareef, DeLiang Wang, Roni Yagel |
IEEE Trans. Medical Imaging | 2 |
| 1999 | Range image segmentation using a relaxation oscillator networkabstractA locally excitatory globally inhibitory oscillator network (LEGION) is constructed and applied to range image segmentation, where each oscillator has excitatory lateral connections to the oscillators in its local neighborhood as well as a connection with a global inhibitor. A feature vector, consisting of depth, surface normal, and mean and Gaussian curvatures, is associated with each oscillator and is estimated from local windows at its corresponding pixel location. A context-sensitive method is applied in order to obtain more reliable and accurate estimations. The lateral connection between two oscillators is established based on a similarity measure of their feature vectors. The emergent behavior of the LEGION network gives rise to segmentation. Due to the flexible representation through phases, our method needs no assumption about the underlying structures in image data and no prior knowledge regarding the number of regions. More importantly, the network is guaranteed to converge rapidly under general conditions. These unique properties may lead to a real-time approach for range image segmentation in machine perception. Xiuwen Liu 0001, DeLiang Wang |
IEEE Trans. Neural Networks | 2 |
| 1999 | Separation of speech from interfering sounds based on oscillatory correlationabstractA multistage neural model is proposed for an auditory scene analysis task--segregating speech from interfering sound sources. The core of the model is a two-layer oscillator network that performs stream segregation on the basis of oscillatory correlation. In the oscillatory correlation framework, a stream is represented by a population of synchronized relaxation oscillators, each of which corresponds to an auditory feature, and different streams are represented by desynchronized oscillator populations. Lateral connections between oscillators encode harmonicity, and proximity in frequency and time. Prior to the oscillator network are a model of the auditory periphery and a stage in which mid-level auditory representations are formed. The model has been systematically evaluated using a corpus of voiced speech mixed with interfering sounds, and produces improvements in terms of signal-to-noise ratio for every mixture. The performance of our model is compared with other studies on computational auditory scene analysis. A number of issues including biological plausibility and real-time implementation are also discussed. DeLiang Wang, Guy J. Brown |
IEEE Trans. Neural Networks | 1 |
| 1998 | Oriented Statistical Nonlinear Smoothing FilterabstractThis paper presents a nonlinear smoothing method which is based on an orientation-sensitive probability measure. By incorporating geometrical constraints through the coupling structure, we obtain a robust nonlinear smoothing algorithm. Even when noise is substantial the proposed smoothing algorithm can still preserve salient boundaries. Compared with anisotropic diffusive approaches, the proposed nonlinear algorithm not only performs better in preserving boundaries but also has a non-uniform stable state, whereby reliable results are available within a fixed number of iterations independent of images. A system using the proposed method and LEGION network has been developed and applied in noisy image segmentation and hydrographic feature extraction from digital ortho-photo quadrangles. Experimental results using synthetic and real images are provided. Xiuwen Liu 0001, DeLiang Wang, J. Raul Ramirez |
ICIP (2) | 2 |
| 1998 | Perceiving without Learning: From Spirals to Inside/Outside Relations
Ke Chen 0001, DeLiang Wang |
NIPS | 2 |
| 1998 | Fast numerical integration of relaxation oscillator networks based on singular limit solutionsabstractAbstract-Relaxation oscillations exhibiting more than one time scale arise naturally from many physical systems. When relaxation oscillators are coupled in a way that resembles chemical synapses, we propose a fast method to numerically integrate such networks. The numerical technique, called the singular limit method, is derived from analysis of relaxation oscillations in the singular limit. In such limit, system evolution gives rise to time instants at which fast dynamics takes place and intervals between them during which slow dynamics takes place. A full description of the method is given for a locally excitatory globally inhibitory oscillator network (LEGION), where fast dynamics, characterized by jumping which leads to dramatic phase shifts, is captured in this method by iterative operation and slow dynamics is entirely solved. The singular limit method is evaluated by computer experiments, and it produces remarkable speedup compared to other methods of integrating these systems. The speedup makes it possible to simulate large-scale oscillator networks. Paul S. Linsay, DeLiang Wang |
IEEE Trans. Neural Networks | 2 |
| 1997 | Image Segmentation Based on Oscillatory CorrelationabstractWe study the image segmentation on the basis of locally excitatory, globally inhibitory oscillator networks (LEGION), whereby the phases of oscillators encode the binding of pixels. We introduce a lateral potential for each oscillators so that only oscillators with strong connections from their neighborhood can develop high potentials. Based on the concept of the lateral potential, a solution to remove noisy regions in an image is proposed for LEGION, so that it suppresses the oscillators corresponding to noisy regions but without affecting those corresponding to major regions. We show that the resulting oscillator network separates an image into several major regions, plus a background consisting of all noisy regions, and we illustrate network properties by computer stimulation. The network exhibits a natural capacity in segmenting images. The oscillatory dynamics leads to a computer algorithm, which is applied successfully to segmenting real gray-level images. A number of issues regarding biological plausibility and perceptual organization are discussed. We argue that LEGION provides a novel and effective framework for image segmentation and figure-ground segregation. DeLiang Wang, David Terman |
Neural Comput. | 1 |
| 1997 | Modelling the perceptual segregation of double vowels with a network of neural oscillators
Guy J. Brown, DeLiang Wang |
Neural Networks | 2 |
| 1996 | On Temporal Generalization of Simple Recurrent Networks
DeLiang Wang, Stanley C. Ahalt |
Neural Networks | 1 |
| 1996 | Synchronization and desynchronization in a network of locally coupled Wilson-Cowan oscillatorsabstractA network of Wilson-Cowan (WC) oscillators is constructed, and its emergent properties of synchronization and desynchronization are investigated by both computer simulation and formal analysis. The network is a 2D matrix, where each oscillator is coupled only to its neighbors. We show analytically that a chain of locally coupled oscillators (the piecewise linear approximation to the WC oscillator) synchronizes, and we present a technique to rapidly entrain finite numbers of oscillators. The coupling strengths change on a fast time scale based on a Hebbian rule. A global separator is introduced which receives input from and sends feedback to each oscillator in the matrix. The global separator is used to desynchronize different oscillator groups. Unlike many other models, the properties of this network emerge from local connections that preserve spatial relationships among components and are critical for encoding Gestalt principles of feature grouping. The ability to synchronize and desynchronize oscillator groups within this network offers a promising approach for pattern segmentation and figure/ground segregation based on oscillatory correlation. Shannon R. Campbell, DeLiang Wang |
IEEE Trans. Neural Networks | 2 |
| 1996 | Incremental learning of complex temporal patternsabstractA neural model for temporal pattern generation is used and analyzed for training with multiple complex sequences in a sequential manner. The network exhibits some degree of interference when new sequences are acquired. It is proven that the model is capable of incrementally learning a finite number of complex sequences. The model is then evaluated with a large set of highly correlated sequences. While the number of intact sequences increases linearly with the number of previously acquired sequences, the amount of retraining due to interference appears to be independent of the size of existing memory. The model is extended to include a chunking network which detects repeated subsequences between and within sequences. The chunking mechanism substantially reduces the amount of retraining in sequential training. Thus, the network investigated here constitutes an effective sequential memory. Various aspects of such a memory are discussed. DeLiang Wang, Budi Yuwono |
IEEE Trans. Neural Networks | 1 |
| 1995 | Emergent synchrony in locally coupled neural oscillatorsabstractThe discovery of long range synchronous oscillations in the visual cortex has triggered much interest in understanding the underlying neural mechanisms and in exploring possible applications of neural oscillations. Many neural models thus proposed end up relying on global connections, leading to the question of whether lateral connections alone can produce remote synchronization. With a formulation different from frequently used phase models, we find that locally coupled neural oscillators can yield global synchrony. The model employs a previously suggested mechanism that the efficacy of the connections is allowed to change on a fast time scale. Based on the known connectivity of the visual cortex, the model outputs closely resemble the experimental findings. Furthermore, we illustrate the potential of locally connected oscillator networks in perceptual grouping and pattern segmentation, which seems missing in globally connected ones. DeLiang Wang |
IEEE Trans. Neural Networks | 1 |
| 1995 | Locally excitatory globally inhibitory oscillator networksabstractA novel class of locally excitatory, globally inhibitory oscillator networks (LEGION) is proposed and investigated. The model of each oscillator corresponds to a standard relaxation oscillator with two time scales. In the network, an oscillator jumping up to its active phase rapidly recruits the oscillators stimulated by the same pattern, while preventing other oscillators from jumping up. Computer simulations demonstrate that the network rapidly achieves both synchronization within blocks of oscillators that are stimulated by connected regions and desynchronization between different blocks. This model lays a physical foundation for the oscillatory correlation theory of feature binding and may provide an effective computational framework for scene segmentation and figure/ground segregation in real time. DeLiang Wang, David Terman |
IEEE Trans. Neural Networks | 1 |
| 1995 | Anticipation-based temporal pattern generationabstractA neural network model of complex temporal pattern generation is proposed and investigated analytically and by computer simulation. Temporal pattern generation is based on recognition of the contexts of individual components. Based on its acquired experience, the model actively yields system anticipation, which then compares with the actual input flow. A mismatch triggers self-organization of context learning, which ultimately leads to resolving various ambiguities in producing complex temporal patterns. The architecture of the model incorporates a short term memory for building associations between remote components and recurrent connections for self-organization and component generation in a temporal pattern. Synaptic modification is based on a one-shot normalized Hebbian rule, which is shown to exhibit temporal masking. The major conclusion, namely the network model which can learn to generate any complex temporal pattern, is established analytically. An estimate on the efficiency of the training algorithm is provided. Multiple temporal patterns can be incrementally acquired by the system, exhibiting a form of retroactive interference. Neural and cognitive plausibility of the model is discussed.> DeLiang Wang, Budi Yuwono |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1994 | An oscillation model of auditory stream segregationabstractAuditory segmentation is critical for complex auditory pattern processing. We present a neural network framework for auditory pattern segmentation. The network is a laterally coupled two-dimensional neural oscillators with a global inhibitor. One dimension represents time and another one represents frequency. We show that this architecture can in real time group auditory features into a segment by phase synchrony and segregate different segments by desynchronization. The network demonstrates the phenomenon that auditory stream segregation critically depends on the rate of presentation. The neuroplausibility of the model is discussed. DeLiang Wang |
ICPR (3) | 1 |
| 1994 | Synchrony and Desynchrony in Neural Oscillator NetworksabstractAn novel class of locally excitatory, globally inhibitory oscillator networks is proposed. The model of each oscillator corresponds to a standard relaxation oscillator with two time scales. The network exhibits a mechanism of selective gating, whereby an oscillator jumping up to its active phase rapidly recruits the oscillators stimulated by the same pattern, while preventing others from jumping up. We show analytically that with the selective gating mechanism the network rapidly achieves both synchronization within blocks of oscillators that are stimulated by connected regions and desynchronization between different blocks. Computer simulations demonstrate the network's promising ability for segmenting multiple input patterns in real time. This model lays a physical foundation for the oscillatory correlation theory of feature binding, and may provide an effective computational framework for scene segmentation and figure/ground segregation. DeLiang Wang, David Terman |
NIPS | 1 |
| 1993 | Timing and chunking in processing temporal orderabstractA computational framework of learning, recognition and reproduction of temporal sequences are provided, based on an interference theory of forgetting in short-term memory (STM), modelled as a network of neural units with mutual inhibition. The STM model provides information for recognition and reproduction of arbitrary temporal sequences. Sequences are acquired by a new learning rule, the attentional learning rule, which combines Hebbian learning and a normalization rule with sequential system activation. Acquired sequences can be recognized without being affected by speed of presentation or certain distortions in symbol form. Different layers of the STM model can be naturally constructed in a feedforward manner to recognize hierarchical sequences, significantly expanding the model's capability in a way similar to human information chunking. A model of sequence reproduction is presented that consists of two reciprocally connected networks, one of which behaves as a sequence recognizer. Reproduction of complex sequences can maintain interval lengths of sequence components, and vary the overall speed. A mechanism of degree self-organization based on a global inhibitor is proposed for the model to learn required context lengths in order to disambiguate associations in complex sequence reproduction. Certain implications of the model are discussed at the end of the paper.> DeLiang Wang, Michael A. Arbib |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1990 | Mechanisms of pattern discrimination in the toad's visual systemabstractThe authors propose that the toad discriminates visual objects based on temporal responses and that discrimination is reflected in different average neuronal firing rates at some higher visual center, hypothetically the anterior thalamus. This theory is developed through a large-scale neuronal simulation which includes retina, tectum, and anterior thalamus. The neural model based on this theory predicts that retinal R2 cells play a primary role in the discrimination via tectal small pear cells (SP) and R3 cells refine the feature analysis by inhibition. The simulation demonstrates that the retinal response to the trailing edge of a stimulus is as crucial for pattern discrimination as the response to the leading edge. The new dishabituation hierarchies are predicted by this model by shrinking stimulus size and reversing contrast DeLiang Wang, Michael A. Arbib |
IJCNN | 1 |
| 1990 | Pattern Segmentation in Associative MemoryabstractThe goal of this paper is to show how to modify associative memory such that it can discriminate several stored patterns in a composite input and represent them simultaneously. Segmention of patterns takes place in the temporal domain, components of one pattern becoming temporally correlated with each other and anticorrelated with the components of all other patterns. Correlations are created naturally by the usual associative connections. In our simulations, temporal patterns take the form of oscillatory bursts of activity. Model oscillators consist of pairs of local cell populations connected appropriately. Transition of activity from one pattern to another is induced by delayed self-inhibition or simply by noise. DeLiang Wang, Joachim M. Buhmann, Christoph von der Malsburg |
Neural Comput. | 1 |
| 1988 | Three neural models which process temporal information
DeLiang Wang, Irwin King |
Neural Networks | 1 |