EDBT 2026 Demo / reviewers in the wild / expert
Yong Xu 0004
dblp:07/4630-4
· DBLP profile ↗
60ranked-venue papers
12as first author
27since 2021 · last 2025
0000-0003-4944-6890ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 9 first-author · 25 since 2021Artificial intelligence and machine learning · 33 · 9 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR ModelsabstractRecent advancements in integrating Large Language Models (LLM) with automatic speech recognition (ASR) have performed remarkably in general domains. While supervised fine-tuning (SFT) of all model parameters is often employed to adapt pre-trained LLM-based ASR models to specific domains, it imposes high computational costs and notably reduces their performance in general domains. In this paper, we propose a novel parameter-efficient multi-domain fine-tuning method for adapting pre-trained LLM-based ASR models to multi-accent domains without catastrophic forgetting named HDMoLE, which leverages hierarchical routing and dynamic thresholds based on combining low-rank adaptation (LoRA) with the mixture of experts (MoE) and can be generalized to any linear layer. Hierarchical routing establishes a clear correspondence between LoRA experts and accent domains, improving cross-domain collaboration among the LoRA experts. Unlike the static Top-K strategy for activating LoRA experts, dynamic thresholds can adaptively activate varying numbers of LoRA experts at each MoE layer. Experiments on the multi-accent and standard Mandarin datasets demonstrate the efficacy of HDMoLE. Applying HDMoLE to an LLM-based ASR model projector module achieves similar performance to full fine-tuning in the target multi-accent domains while using only 9.6% of the trainable parameters required for full fine-tuning and minimal degradation in the source general domain. Bingshen Mu, Qijie Shao, Yong Xu 0004, Lei Xie 0001 |
ICASSP | 4 |
| 2025 | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2025 | VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
Jianheng Zhuo, Yifan Yang 0005, Yiwen Shao, Yong Xu 0004, Dong Yu 0001, Kai Yu 0004, Xie Chen 0001 |
INTERSPEECH | 4 |
| 2024 | SPATIALCODEC: Neural Spatial Speech CodingabstractIn this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We propose a neural spatial audio coding framework that achieves a high compression ratio, leveraging single-channel neural sub-band codec and SpatialCodec. Our approach encompasses two phases: (i) a neural sub-band codec is designed to encode the reference channel with low bit rates, and (ii), a SpatialCodec captures relative spatial information for accurate multi-channel reconstruction at the decoder end. In addition, we also propose novel evaluation metrics to assess the spatial cue preservation: (i) spatial similarity, which calculates cosine similarity on a spatially intuitive beamspace, and (ii), beamformed audio quality. Our system shows superior spatial performance compared with high bitrate baselines and black-box neural architecture. Demos are available at https://xzwy.github.io/SpatialCodecDemo. Codes and models are available at https://github.com/XZWY/SpatialCodec. Zhongweiyang Xu, Yong Xu 0004, Vinay Kothapally, Heming Wang, Muqiao Yang, Dong Yu 0001 |
ICASSP | 2 |
| 2024 | uSee: Unified Speech Enhancement And Editing with Conditional Diffusion ModelsabstractSpeech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional diffusion models to handle various tasks at the same time in a generative manner. Specifically, by providing multiple types of conditions including self-supervised learning embeddings and proper text prompts to the score-based diffusion model, we can enable controllable generation of the unified speech enhancement and editing model to perform corresponding actions on the source speech. Our experiments show that our proposed uSee model can achieve superior performance in both speech denoising and dereverberation compared to other related generative speech enhancement models, and can perform speech editing given desired environmental sound text description, signal-to-noise ratios (SNR), and room impulse responses (RIR). Demos of the generated speech are available at https://muqiaoy.github.io/usee. Muqiao Yang, Yong Xu 0004, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu 0001 |
ICASSP | 3 |
| 2024 | LibriheavyMix: A 20, 000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
Zengrui Jin, Yifan Yang 0005, Mohan Shi, Wei Kang 0006, Xiaoyu Yang 0005, Zengwei Yao, Liyong Guo, Lingwei Meng, Long Lin, Yong Xu 0004, Shixiong Zhang 0001, Daniel Povey |
INTERSPEECH | 11 |
| 2024 | Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2024 | Spatialemb: Extract and Encode Spatial Information for 1-Stage Multi-Channel Multi-Speaker ASR on Arbitrary Microphone ArraysabstractSpatial information is a critical clue for multi-channel multispeaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation stage, followed by standard single-channel ASR on the separated speech. This approach results in an inefficient, lengthy pipeline and sub-optimal ASR performance due to the accumulated errors from preprocessing modules. Furthermore, most spatial feature extraction methods depend on the knowledge of speaker positions and microphone topology, making the systems reliant on specific settings and challenging to adapt to new equipment. In this work, we propose a solution to these issues with a lightweight embedding module named SpatialEmb, which extracts and encodes spatial information directly for the ASR model, supporting both fixed and arbitrary microphone topology. We conduct comprehensive experiments on AliMeeting, a real meeting corpus, to determine the optimal model design for SpatialEmb in terms of both performance and efficiency. Our best model trained with 105 hours Train-Ali-far achieves 17.04% and 20.32% character error rates (CER) on the Eval and Test sets, establishing a new state-of-the-art result with the same training data. Yiwen Shao, Yong Xu 0004, Sanjeev Khudanpur, Dong Yu 0001 |
SLT | 2 |
| 2024 | Advancing Multi-Talker ASR Performance With Large Language ModelsabstractRecognizing overlapping speech from multiple speakers in conversational scenarios is one of the most challenging problem for automatic speech recognition (ASR). Serialized output training (SOT) is a classic method to address multi-talker ASR, with the idea of concatenating transcriptions from multiple speakers according to the emission times of their speech for training. However, SOT-style transcriptions, derived from concatenating multiple related utterances in a conversation, depend significantly on modeling long contexts. Therefore, compared to traditional methods that primarily emphasize encoder performance in attention-based encoderdecoder (AED) architectures, a novel approach utilizing large language models (LLMs) that leverages the capabilities of pre-trained decoders may be better suited for such complex and challenging scenarios. In this paper, we propose an LLM-based SOT approach for multi-talker ASR, leveraging pre-trained speech encoder and LLM, fine-tuning them on multi-talker dataset using appropriate strategies. Experimental results demonstrate that our approach surpasses traditional AED-based methods on the simulated dataset LibriMix and achieves state-of-the-art performance on the evaluation set of the real-world dataset AMI, outperforming the AED model trained with 1000 times more supervised data in previous works. Mohan Shi, Zengrui Jin, Yaoxun Xu, Yong Xu 0004, Shixiong Zhang 0001, Yiwen Shao, Dong Yu 0001 |
SLT | 4 |
| 2024 | Labelled Non-Zero Diffusion Particle Flow SMC-PHD Filtering for Multi-Speaker TrackingabstractParticle flow (PF) is a method originally proposed for single target tracking, and used recently to address the weight degeneracy problem of the sequential Monte Carlo probability hypothesis density (SMC-PHD) filter for audio-visual (AV) multi-speaker tracking, where the particle flow is calculated by using only the measurements near the particle, assuming that the target is detected, as in a recent method based on non-zero particle flow (NPF), i.e. the AV-NPF-SMC-PHD filter. This, however, can be problematic when occlusion happens and the occluded speaker may not be detected. To address this issue, we propose a new method where the labels of the particles are estimated using the likelihood function, and the particle flow is calculated in terms of the selected particles with the same labels. As a result, the particles associated with detected speakers and undetected speakers are distinguished based on the particle labels. With this novel method, named as AV-LPF-SMC-PHD, the speaker states can be estimated as the weighted mean of the labelled particles, which is computationally more efficient than using a clustering method as in the AV-NPF-SMC-PHD filter. The proposed algorithm is compared systematically with several baseline tracking methods using the AV16.3, AVDIAR and CLEAR datasets, and is shown to offer improved tracking accuracy with a lower computational cost. Yang Liu 0175, Yong Xu 0004, Peipei Wu, Wenwu Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Neuralecho: Hybrid of Full-Band and Sub-Band Recurrent Neural Network For Acoustic Echo Cancellation and Speech EnhancementabstractThis paper presents a hybrid of full-band and sub-band recurrent neural network (RNN) model, named NeuralEcho, to jointly solve echo and noise suppression. The full-band model part processes the signal’s entire frequency bands as a whole, while the sub-band model part divides the features into sub-bands and processes each sub-band separately. This approach allows the model to capture both the fine-grained local details of the sub-band processing and the global context of the full-band processing. The single-channel model is then generalized to accommodate a range of input channel numbers. Experimental results show that the hybrid model outperforms the conventional full-band models in terms of objective speech quality metrics and speech recognition accuracy. This suggests that the hybrid approach of full-band and sub-band processing can be a promising direction for future research in the field of speech enhancement. Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
ASRU | 2 |
| 2023 | Deep Neural Mel-Subband Beamformer for in-Car Speech SeparationabstractWhile current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies independently which results in higher computational costs and inference times, making them unsuitable for real-world use. In this paper, we propose DL-based mel-subband spatio-temporal beamformer to perform speech separation in a car environment with reduced computation cost and inference time. As opposed to conventional subband (SB) approaches, our framework uses a mel-scale based subband selection strategy which ensures a fine-grained processing for lower frequencies where most speech formant structure is present, and coarse-grained processing for higher frequencies. In a recursive way, robust frame-level beamforming weights are determined for each speaker location/zone in a car from the estimated subband speech and noise covariance matrices. Furthermore, proposed framework also estimates and suppresses any echoes from the loudspeaker(s) by using the echo reference signals. We compare the performance of our proposed framework to several NB, SB, and full-band (FB) processing techniques in terms of speech quality and recognition metrics. Based on experimental evaluations on simulated and real-world recordings, we find that our proposed framework achieves better separation performance over all SB and FB approaches and achieves performance closer to NB processing techniques while requiring lower computing cost. Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
ICASSP | 2 |
| 2023 | Zoneformer: On-device Neural Beamformer For In-car Multi-zone Speech Separation, Enhancement and Echo Cancellation
Yong Xu 0004, Vinay Kothapally, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
INTERSPEECH | 1 |
| 2022 | Audio-Visual Tracking of Multiple Speakers Via a PMBM FilterabstractAudio-visual tracking of multiple speakers requires to estimate the state (e.g. velocity and location) of each speaker by leveraging the information of both audio and visual modalities. Estimating the number of speakers and their states jointly remains a challenging problem. We propose an Audio-Visual Possion Multi-Bernoulli Mixture Filter (AV-PMBM) that can not only predict the number of speakers but also give accurate estimation of their states. We also propose a novel sound source localization technique based on DOA information and a deep learning based object detector to provide reliable audio measurements for the AV tracker. To our knowledge, this represents the first attempt using PMBM for multi-speaker tracking with audio visual modalities. Experiments on the AV16.3 dataset demonstrate that AV-PMBM achieves state-of-the-art performance in optimal sub-pattern assignment (OSPA). Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Yong Xu 0004, Lyudmila Mihaylova, Simon J. Godsill, Wenwu Wang 0001 |
ICASSP | 4 |
| 2022 | Joint Neural AEC and Beamforming with Double-Talk DetectionabstractAcoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback.However, nonlinear distortions induced by audio devices, background noise, reverberation, and double-talk reduce the efficiency of conventional AEC systems.Several hybrid AEC models were proposed to address this, which use deep learning models to suppress residual echo from standard adaptive filtering.This paper proposes deep learning-based joint AEC and beamforming model (JAECBF) building on our previous self-attentive recurrent neural network (RNN) beamformer.The proposed network consists of two modules: (i) multi-channel neural-AEC, and (ii) joint AEC-RNN beamformer with a double-talk detection (DTD) that computes time-frequency (T-F) beamforming weights.We train the proposed model in an end-to-end approach to eliminate background noise and echoes from far-end audio devices, which include nonlinear distortions.From experimental evaluations, we find the proposed network outperforms other multi-channel AEC and denoising systems in terms of speech recognition rate and overall speech quality. Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2022 | Audio Visual Multi-Speaker Tracking with Improved GCF and PMBM Filter
Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Shidrokh Goudarzi, Haohe Liu, Yong Xu 0004, Wenwu Wang 0001 |
INTERSPEECH | 6 |
| 2022 | EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of SpeakersabstractIn this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neural diarization (EEND) models, speaker counting with encoder-decoder based attractors (EDA), and speech separation using Conv-TasNet. In addition, we propose a multiple$1 \times 1$convolutional layer architecture for estimating the separation masks corresponding to a flexible number of speakers and a fusion technique for refining the separated speech signal with obtained speaker diarization information to improve the joint framework. Experiments using the LibriMix dataset show that our proposed method outperforms the single-task baselines in both diarization and separation metrics for fixed and flexible numbers of speakers and improves speaker counting performance for flexible numbers of speakers. All materials will be open-sourced and reproducible in ESPnet toolkit11https://github.com/espnet/espnet. Soumi Maiti, Yushi Ueda, Shinji Watanabe 0001, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004 |
SLT | 7 |
| 2021 | Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source LocalizationabstractThis paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end (E2E) neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the azimuth angle of the sources with respect to the microphone array is defined as a latent variable. This angle controls the quality of separation, which in turn determines the ASR performance. All three functionalities of D-ASR: localization, separation, and recognition are connected as a single differentiable neural network and trained solely based on ASR error minimization objectives. The advantages of D-ASR over existing methods are threefold: (1) it provides explicit speaker locations, (2) it improves the explainability factor, and (3) it achieves better ASR performance as the process is more streamlined. In addition, D-ASR does not require explicit direction of arrival (DOA) supervision like existing data-driven localization models, which makes it more appropriate for realistic data. For the case of two source mixtures, D-ASR achieves an average DOA prediction error of less than three degrees. It also outperforms a strong far-field multi-speaker end-to-end system in both separation quality and ASR performance. Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
ICASSP | 5 |
| 2021 | ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech SeparationabstractSpeech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech recognition (ASR) systems. The conventional mask-based minimum variance distortionless response (MVDR) beamformer can be used to minimize the distortion, but comes with high level of residual noise. Furthermore, the matrix operations (e.g., matrix inversion) involved in the conventional MVDR solution are sometimes numerically unstable when jointly trained with neural networks. In this paper, we propose a novel all deep learning MVDR framework, where the matrix inversion and eigenvalue decomposition are replaced by two recurrent neural networks (RNNs), to resolve both issues at the same time. The proposed method can greatly reduce the residual noise while keeping the target speech undistorted by leveraging on the RNN-predicted frame-wise beamforming weights. The system is evaluated on a Mandarin audio-visual corpus and compared against several state-of-the-art (SOTA) speech separation systems. Experimental results demonstrate the superiority of the proposed method across several objective metrics and ASR accuracy. Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001 |
ICASSP | 2 |
| 2021 | MIMO Self-Attentive RNN Beamformer for Multi-Speaker Speech SeparationabstractRecently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion and eigenvalue decomposition with two RNNs.In this work, we present a self-attentive RNN beamformer to further improve our previous RNN-based beamformer by leveraging on the powerful modeling capability of self-attention.Temporal-spatial self-attention module is proposed to better learn the beamforming weights from the speech and noise spatial covariance matrices.The temporal self-attention module could help RNN to learn global statistics of covariance matrices.The spatial self-attention module is designed to attend on the cross-channel correlation in the covariance matrices.Furthermore, a multi-channel input with multi-speaker directional features and multi-speaker speech separation outputs (MIMO) model is developed to improve the inference efficiency.The evaluations demonstrate that our proposed MIMO self-attentive RNN beamformer improves both the automatic speech recognition (ASR) accuracy and the perceptual estimation of speech quality (PESQ) against prior arts. Xiyun Li, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Jiaming Xu 0001, Bo Xu 0002, Dong Yu 0001 |
Interspeech | 2 |
| 2021 | TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech DereverberationabstractIn this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach on the deep neural network (DNN) for environment-aware speech dereverberation, which can adaptively attend to the contextual information. More specifically, a FullBand based Temporal Attention approach (FTA) is proposed, which models the correlations between the fullband information of the context frames. In addition, considering the difference between the attenuation of high frequency bands and low frequency bands (high frequency bands attenuate faster than low frequency bands) in the room impulse response (RIR), we also propose a SubBand based Temporal Attention approach (STA). In order to guide the network to be more aware of the reverberant environments, we jointly optimize the dereverberation network and the reverberation time (RT60) estimator in a multi-task manner. Our experimental results indicate that the proposed method outperforms our previously proposed reverberation-time-aware DNN and the learned attention weights are fully physical consistent. We also report a preliminary yet promising dereverberation and recognition experiment on real test data. Helin Wang, Bo Wu 0011, Lianwu Chen, Meng Yu 0003, Jianwei Yu 0001, Yong Xu 0004, Shixiong Zhang 0001, Chao Weng, Dan Su 0002, Dong Yu 0001 |
Interspeech | 6 |
| 2021 | Generalized Spatio-Temporal RNN Beamformer for Target Speech SeparationabstractAlthough the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal recurrent neural network based beamformer (RNN-BF) for target speech separation. This new beamforming framework directly learns the beamforming weights from the estimated speech and noise spatial covariance matrices. Leveraging on the temporal modeling capability of RNNs, the RNN-BF could automatically accumulate the statistics of the speech and noise covariance matrices to learn the frame-level beamforming weights in a recursive way. An RNN-based generalized eigenvalue (RNN-GEV) beamformer and a more generalized RNN beamformer (GRNN-BF) are proposed. We further improve the RNN-GEV and the GRNN-BF by using layer normalization to replace the commonly used mask normalization on the covariance matrices. The proposed GRNN-BF obtains better performance against prior arts in terms of speech quality (PESQ), speech-to-noise ratio (SNR) and word error rate (WER). Yong Xu 0004, Zhuohuang Zhang, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
Interspeech | 1 |
| 2021 | MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality AssessmentabstractThe objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion score (MOS) tests. Non-intrusive speech quality assessment has attracted much attention recently due to the lack of access to clean reference signals for objective evaluations in real scenarios. In this paper, we propose a novel non-intrusive speech quality measurement model, MetricNet, which leverages label distribution learning and joint speech reconstruction learning to achieve significantly improved performance compared to the existing non-intrusive speech quality measurement models. We demonstrate that the proposed approach yields promisingly high correlation to the intrusive objective evaluation of speech quality on clean, noisy and processed speech data. Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
Interspeech | 3 |
| 2021 | Neural Mask based Multi-channel Convolutional Beamforming for Joint Dereverberation, Echo Cancellation and DenoisingabstractThis paper proposes a new joint optimization framework for simultaneous dereverberation, acoustic echo cancellation, and denoising, which is motivated by the recently proposed con-volutional beamformer for simultaneous denoising and dereverberation. Using the echo aware mask based beamforming framework, the proposed algorithm could effectively deal with double-talk case and local inference, etc. The evaluations based on ERLE for echo only, and PESQ for double-talk demonstrate that the proposed algorithm could significantly improve the performance. Meng Yu 0003, Yong Xu 0004, Chao Weng, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001 |
SLT | 3 |
| 2021 | WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and DereverberationabstractThis paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted Power minimization Distortionless response (WPD) beamformer can perform separation and dereverberation simultaneously, the noise cancellation component still has the potential to progress. We propose an improved neural WPD beamformer called "WPD++" by an enhanced beamforming module in the conventional WPD and a multi-objective loss function for the joint training. The beamforming module is improved by utilizing the spatio-temporal correlation. A multi-objective loss, including the complex spectra domain scale-invariant signal-to-noise ratio (C-Si-SNR) and the magnitude domain mean square error (Mag-MSE), is properly designed to make multiple constraints on the enhanced speech and the desired power of the dry clean signal. Joint training is conducted to optimize the complex-valued mask estimator and the WPD++ beamformer in an end-to-end way. The results show that the proposed WPD++ outperforms several state-of-the-art beamformers on the enhanced speech quality and word error rate (WER) of ASR. Zhaoheng Ni, Yong Xu 0004, Meng Yu 0003, Bo Wu 0011, Shixiong Zhang 0001, Dong Yu 0001, Michael I. Mandel |
SLT | 2 |
| 2021 | An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and SeparationabstractSpeech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance. Daniel Michelsanti, Zheng-Hua Tan, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Multi-Channel Multi-Frame ADL-MVDR for Target Speech SeparationabstractMany purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems. Minimum variance distortionless response (MVDR) filters are often adopted to remove nonlinear distortions, however, conventional neural mask-based MVDR systems still result in relatively high levels of residual noise. Moreover, the matrix inverse involved in the MVDR solution is sometimes numerically unstable during joint training with neural networks. In this study, we propose a multi-channel multi-frame (MCMF) all deep learning (ADL)-MVDR approach for target speech separation, which extends our preliminary multi-channel ADL-MVDR approach. The proposed MCMF ADL-MVDR system addresses linear and nonlinear distortions. Spatio-temporal cross correlations are also fully utilized in the proposed approach. The proposed systems are evaluated using a Mandarin audio-visual corpus and are compared with several state-of-the-art approaches. Experimental results demonstrate the superiority of our proposed systems under different scenarios and across several objective evaluation metrics, including ASR performance. Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Donald S. Williamson, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Self-Supervised Learning for Audio-Visual Speaker DiarizationabstractSpeaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propose a self-supervised audio-video synchronization learning method to address the problem of speaker diarization without massive labeling effort. We improve the previous approaches by introducing two new loss functions: the dynamic triplet loss and the multinomial loss. We test them on a real-world human-computer interaction system and the results show our best model yields a remarkable gain of +8% F1-scores as well as diarization error rate reduction. Finally, we introduce a new large scale audio-video corpus designed to fill the vacancy of audio-video dataset in Chinese. Yifan Ding 0002, Yong Xu 0004, Shixiong Zhang 0001, Yahuan Cong, Liqiang Wang 0001 |
ICASSP | 2 |
| 2020 | Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature LearningabstractHand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In this work, we propose an integrated architecture for learning spatial features directly from the multi-channel speech waveforms within an end-to-end speech separation framework. In this architecture, time-domain filters spanning signal channels are trained to perform adaptive spatial filtering. These filters are implemented by a 2d convolution (conv2d) layer and their parameters are optimized using a speech separation objective function in a purely data-driven fashion. Furthermore, inspired by the IPD formulation, we design a conv2d kernel to compute the inter-channel convolution differences (ICDs), which are expected to provide the spatial cues that help to distinguish the directional sources. Evaluation results on simulated multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed ICD based MCSS model improves the overall signal-to-distortion ratio by 10.4% over the IPD based MCSS model. Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001 |
ICASSP | 4 |
| 2020 | Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition ObjectivesabstractTarget speech extraction is a specific case of source separation where an auxiliary information like the location or some pre-saved anchor speech examples of the target speaker is used to resolve the permutation ambiguity. Traditionally such systems are optimized based on signal reconstruction objectives. Recently end-to-end automatic speech recognition (ASR) methods have enabled to optimize source separation systems with only the transcription based objective. This paper proposes a method to jointly optimize a location guided target speech extraction module along with a speech recognition module only with ASR error minimization criteria. Experimental comparisons with corresponding conventional pipeline systems verify that this task can be realized by end-to-end ASR training objectives without using parallel clean data. We show promising target speech recognition results in mixtures of two speakers and noise, and discuss interesting properties of the proposed system in terms of speech enhancement/separation objectives and word error rates. Finally, we design a system that can take both location and anchor speech as input at the same time and show that the performance can be further improved. Aswin Shanmugam Subramanian, Chao Weng, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004, Shinji Watanabe 0001, Dong Yu 0001 |
ICASSP | 5 |
| 2020 | Neural Spatio-Temporal Beamformer for Target Speech SeparationabstractPurely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR).On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability.In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement.Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the intermicrophone correlation that is already utilized in prior arts.Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN.The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts. Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Chao Weng, Dong Yu 0001 |
INTERSPEECH | 1 |
| 2020 | Audio-Visual Multi-Channel Recognition of Overlapped SpeechabstractAutomatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption, this paper presents an audio-visual multi-channel overlapped speech recognition system featuring tightly integrated separation front-end and recognition back-end. A series of audio-visual multi-channel speech separation front-end components based on \textit{TF masking}, \textit{filter\&sum} and \textit{mask-based MVDR} beamforming approaches were developed. To reduce the error cost mismatch between the separation and recognition components, they were jointly fine-tuned using the connectionist temporal classification (CTC) loss function, or a multi-task criterion interpolation with scale-invariant signal to noise ratio (Si-SNR) error cost. Experiments suggest that the proposed multi-channel AVSR system outperforms the baseline audio-only ASR system by up to 6.81\% (26.83\% relative) and 22.22\% (56.87\% relative) absolute word error rate (WER) reduction on overlapped speech constructed using either simulation or replaying of the lipreading sentence 2 (LRS2) dataset respectively. Jianwei Yu 0001, Bo Wu 0011, Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Dong Yu 0001, Xunying Liu, Helen M. Meng |
INTERSPEECH | 6 |
| 2020 | Sound Event Detection of Weakly Labelled Data With CNN-Transformer and Automatic Threshold OptimizationabstractSound event detection (SED) is a task to detect sound events in an audio recording. One challenge of the SED task is that many datasets such as the Detection and Classification of Acoustic Scenes and Events (DCASE) datasets are weakly labelled. That is, there are only audio tags for each audio clip without the onset and offset times of sound events. We compare segment-wise and clip-wise training for SED that is lacking in previous works. We propose a convolutional neural network transformer (CNN-Transfomer) for audio tagging and SED, and show that CNN-Transformer performs similarly to a convolutional recurrent neural network (CRNN). Another challenge of SED is that thresholds are required for detecting sound events. Previous works set thresholds empirically, and are not an optimal approaches. To solve this problem, we propose an automatic threshold optimization method. The first stage is to optimize the system with respect to metrics that do not depend on thresholds, such as mean average precision (mAP). The second stage is to optimize the thresholds with respect to metrics that depends on those thresholds. Our proposed automatic threshold optimization system achieves a state-of-the-art audio tagging F1 of 0.646, outperforming that without threshold optimization of 0.629, and a sound event detection F1 of 0.584, outperforming that without threshold optimization of 0.564. Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Time Domain Audio Visual Speech SeparationabstractAudio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker extraction from monaural mixtures. The architecture generalizes the previous TasNet (time-domain speech separation network) to enable multi-modal learning and at meanwhile it extends the classical audio-visual speech separation from frequency-domain to time-domain. The main components of proposed architecture include an audio encoder, a video encoder that extracts lip embedding from video streams, a multi-modal separation network and an audio decoder. Experiments on simulated mixtures based on recently released LRS2 dataset show that our method can bring 3dB+ and 4dB+ Si-SNR improvements on two- and three-speaker cases respectively, compared to audio-only TasNet and frequency-domain audio-visual networks. Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001 |
ASRU | 2 |
| 2019 | An Attention-based Neural Network Approach for Single Channel Speech EnhancementabstractThis paper proposes an attention-based neural network approach for single channel speech enhancement. Our work is inspired by the recent success of attention models in sequence-to-sequence learning. It is intuitive to use attention mechanism in speech enhancement as humans are able to focus on the important speech components in an audio stream with "high attention" while perceiving the unimportant region (e.g., noise or interference) in "low attention", and thus adjust the focal point over time. Specifically, taking noisy spectrum as input, our model is composed of an LSTM based encoder, an attention mechanism and a speech generator, resulting in enhanced spectrum. Experiments show that, as compared with OM-LSA and the LSTM baseline, the proposed attention approach can consistently achieve better performance in terms of speech quality (PESQ) and intelligibility (STOI). More promisingly, the attention-based approach has better generalization ability to unseen noise conditions. Changhao Shan, Yong Xu 0004, Sining Sun, Lei Xie 0001 |
ICASSP | 3 |
| 2019 | Acoustic Scene Generation with Conditional SamplernnabstractAcoustic scene generation (ASG) is a task to generate waveforms for acoustic scenes. ASG can be used to generate audio scenes for movies and computer games. Recently, neural networks such as SampleRNN have been used for speech and music generation. However, ASG is more challenging due to its wide variety. In addition, evaluating a generative model is also difficult. In this paper, we propose to use a conditional SampleRNN model to generate acoustic scenes conditioned on the input classes. We also propose objective criteria to evaluate the quality and diversity of the generated samples based on classification accuracy. The experiments on the DCASE 2016 Task 1 acoustic scene data show that with the generated audio samples, a classification accuracy of 65.5% can be achieved compared to samples generated by a random model of 6.7% and samples from real recording of 83.1%. The performance of a classifier trained only on generated samples achieves an accuracy of 51.3%, as opposed to an accuracy of 6.7% with samples generated by a random model. Qiuqiang Kong, Yong Xu 0004, Turab Iqbal, Yin Cao, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 2 |
| 2019 | Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust AsrabstractIn this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covariance matrix for the beamformer. Minimum Variance Distortionless Response (MVDR) beamformer and Generalized Eigenvalue (GEV) beamformer are both investigated under the CRM-based joint training architecture. We also propose a robust mask pooling strategy among multiple channels. A long short-term memory (LSTM) based language model is utilized to re-score hypotheses which further improves the overall performance. We evaluate the proposed methods on CHiME-4 challenge dataset. The CRM based system achieves a relative 10% reduction on word error rate (WER) compared with the IRM based system. Without sequence discriminative training, our best single system already achieves an average WER 2.72% on the test set which is comparable to the state-of-the-art. Yong Xu 0004, Chao Weng, Like Hui, Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 1 |
| 2019 | Single-Channel Signal Separation and Deconvolution with Generative Adversarial NetworksabstractSingle-channel signal separation and deconvolution aims to separate and deconvolve individual sources from a single-channel mixture. Single-channel signal separation and deconvolution is a challenging problem in which no prior knowledge of the mixing filters is available. Both individual sources and mixing filters need to be estimated. In addition, a mixture may contain non-stationary noise which is unseen in the training set. We propose a synthesizing-decomposition (S-D) approach to solve the single-channel separation and deconvolution problem. In synthesizing, a generative model for sources is built using a generative adversarial network (GAN). In decomposition, both mixing filters and sources are optimized to minimize the reconstruction error of the mixture. The proposed S-D approach achieves a peak-to-noise-ratio (PSNR) of 18.9 dB and 15.4 dB in image inpainting and completion, outperforming a baseline convolutional neural network PSNR of 15.3 dB and 12.2 dB, respectively and achieves a PSNR of 13.2 dB in source separation together with deconvolution, outperforming a convolutive non-negative matrix factorization (NMF) baseline of 10.1 dB. Qiuqiang Kong, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Mark D. Plumbley |
IJCAI | 2 |
| 2019 | A Comprehensive Study of Speech Separation: Spectrogram vs Waveform SeparationabstractSpeech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain.Recently, a raw audio waveform separation network (TasNet) is introduced for single-channel data, with achieving high Si-SNR (scale-invariant source-to-noise ratio) and SDR (sourceto-distortion ratio) comparing against the state-of-the-art solution in frequency-domain.In this study, we incorporate effective components of the TasNet into a frequency-domain separation method.We compare both for alternative scenarios.We introduce a solution for directly optimizing the separation criterion in frequency-domain networks.In addition to speech separation objective and subjective measurements, we evaluate the separation performance on a speech recognition task as well.We study the speech separation problem for far-field data (more similar to naturalistic audio streams) and develop multi-channel solutions for both frequency and time-domain separators with utilizing spectral, spatial and speaker location information.For our experiments, we simulated multi-channel spatialized reverberate WSJ0-2mix dataset.Our experimental results show that spectrogram separation can achieve competitive performance with better network design.Multi-channel framework as well is shown to improve the single-channel performance relatively up to +35.5% and +46% in terms of WER and SDR, respectively. Fahimeh Bahmaninezhad, Jian Wu 0027, Rongzhi Gu, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2019 | Neural Spatial Filter: Target Speaker Speech Separation Assisted with Directional Information
Rongzhi Gu, Lianwu Chen, Shixiong Zhang 0001, Jimeng Zheng, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2019 | Improved Speaker-Dependent Separation for CHiME-5 ChallengeabstractThis paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-channel, highly-overlapped conversational speech recognition in a dinner party scenario with reverberations and nonstationary noises.We adopt a speaker-aware training method by using i-vector as the target speaker information for multi-talker speech separation.With only one unified separation model for all speakers, we achieve a 10% absolute improvement in terms of word error rate (WER) over the previous baseline of 80.28% on the development set by leveraging our newly proposed data processing techniques and beamforming approach.With our improved back-end acoustic model, we further reduce WER to 60.15% which surpasses the result of our submitted CHiME-5 challenge system without applying any fusion techniques. Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2019 | Sound Event Detection and Time-Frequency Segmentation from Weakly Labelled DataabstractSound event detection (SED) aims to detect when and recognize what sound events happen in an audio clip. Many supervised SED algorithms rely on strongly labelled data that contains the onset and offset annotations of sound events. However, many audio tagging datasets are weakly labelled, that is, only the presence of the sound events is known, without knowing their onset and offset annotations. In this paper, we propose a time–frequency (T–F) segmentation framework trained on weakly labelled data to tackle the sound event detection and separation problem. In training, a segmentation mapping is applied on a T–F representation, such as log mel spectrogram of an audio clip to obtain T–F segmentation masks of sound events. The T–F segmentation masks can be used for separating the sound events from the background scenes in the T–F domain. Then, a classification mapping is applied on the T–F segmentation masks to estimate the presence probabilities of the sound events. We model the segmentation mapping using a convolutional neural network and the classification mapping using a global weighted rank pooling. In SED, predicted onset and offset times can be obtained from the T–F segmentation masks. As a byproduct, separated waveforms of sound events can be obtained from the T–F segmentation masks. We remixed the DCASE 2018 Task 1 acoustic scene data with the DCASE 2018 Task 2 sound events data. When mixing under 0 dB, the proposed method achieved F1 scores of 0.534, 0.398, and 0.167 in audio tagging, frame-wise SED and event-wise SED, outperforming the fully connected deep neural network baseline of 0.331, 0.237, and 0.120, respectively. In T–F segmentation, we achieved an F1 score of 0.218, where previous methods were not able to do T–F segmentation. Qiuqiang Kong, Yong Xu 0004, Iwona Sobieraj, Wenwu Wang 0001, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Weakly Labelled AudioSet Tagging With Attention Neural NetworksabstractAudio tagging is the task of predicting the presence or absence of sound classes within an audio clip. Previous work in audio tagging focused on relatively small datasets limited to recognizing a small number of sound classes. We investigate audio tagging on AudioSet, which is a dataset consisting of over 2 million audio clips and 527 classes. AudioSet is weakly labelled, in that only the presence or absence of sound classes is known for each clip, whereas the onset and offset times are unknown. To address the weakly labelled audio tagging problem, we propose attention neural networks as a way to attend the most salient parts of an audio clip. We bridge the connection between attention neural networks and multiple instance learning (MIL) methods, and propose decision-level and feature-level attention neural networks for audio tagging. We investigate attention neural networks modeled by different functions, depths, and widths. Experiments on AudioSet show that the feature-level attention neural network achieves a state-of-the-art mean average precision of 0.369, outperforming the best MIL method of 0.317 and Google's deep neural network baseline of 0.314. In addition, we discover that the audio tagging performance on AudioSet-embedding features has a weak correlation with the number of training samples and the quality of labels of each sound class. Qiuqiang Kong, Changsong Yu, Yong Xu 0004, Turab Iqbal, Wenwu Wang 0001, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Large-Scale Weakly Supervised Audio Classification Using Gated Convolutional Neural NetworkabstractIn this paper, we present a gated convolutional neural network and a temporal attention-based localization method for audio classification, which won the 1st place in the large-scale weakly supervised sound event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 challenge. The audio clips in this task, which are extracted from YouTube videos, are manually labelled with one or more audio tags, but without time stamps of the audio events, hence referred to as weakly labelled data. Two subtasks are defined in this challenge including audio tagging and sound event detection using this weakly labelled data. We propose a convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) non-linearity applied on the log Mel spectrogram. In addition, we propose a temporal attention method along the frames to predict the locations of each audio event in a chunk from the weakly labelled data. The performances of our systems were ranked the 1st and the 2nd as a team in these two sub-tasks of DCASE 2017 challenge with F value 55.6% and Equal error 0.73, respectively. Yong Xu 0004, Qiuqiang Kong, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 1 |
| 2018 | Audio Set Classification with Attention Model: A Probabilistic PerspectiveabstractThis paper investigates the Audio Set classification. Audio Set is a large scale weakly labelled dataset (WLD) of audio clips. In WLD only the presence of a label is known, without knowing the happening time of the labels. We propose an attention model to solve this WLD problem and explain the attention model from a novel probabilistic perspective. Each audio clip in Audio Set consists of a collection of features. We call each feature as an instance and the collection as a bag following the terminology in multiple instance learning. In the attention model, each instance in the bag has a trainable probability measure for each class. The classification of the bag is the expectation of the classification output of the instances in the bag with respect to the learned probability measure. Experiments show that the proposed attention model achieves a mAP of 0.327 on Audio Set, outperforming the Google's baseline of 0.314. Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 2 |
| 2018 | A Joint Separation-Classification Model for Sound Event Detection of Weakly Labelled DataabstractSource separation (SS) aims to separate individual sources from an audio recording. Sound event detection (SED) aims to detect sound events from an audio recording. We propose a joint separation-classification (JSC) model trained only on weakly labelled audio data, that is, only the tags of an audio recording are known but the time of the events are unknown. First, we propose a separation mapping from the time-frequency (T-F) representation of an audio to the T-F segmentation masks of the audio events. Second, a classification mapping is built from each T-F segmentation mask to the presence probability of each audio event. In the source separation stage, sources of audio events and time of sound events can be obtained from the T-F segmentation masks. The proposed method achieves an equal error rate (EER) of 0.14 in SED, outperforming deep neural network baseline of 0.29. Source separation SDR of 8.08 dB is obtained by using global weighted rank pooling (GWRP) as probability mapping, out-performing the global max pooling (GMP) based probability mapping giving SDR at 0.03 dB. Source code of our work is published. Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 2 |
| 2018 | Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech SeparationabstractIn this paper, we propose an iterative deep neural network (DNN)-based binaural source separation scheme, for recovering two concurrent speech signals in a room environment. Besides the commonly-used spectral features, the DNN also takes non-linearly wrapped binaural spatial features as input, which are refined iteratively using parameters estimated from the DNN output via a feedback loop. Different DNN structures have been tested, including a classic multilayer perception regression architecture as well as a new hybrid network with both convolutional and densely-connected layers. Objective evaluations in terms of PESQ and STOI showed consistent improvement over baseline methods using traditional binaural features, especially when the hybrid DNN architecture was employed. In addition, our proposed scheme is robust to mismatches between the training and testing data. Qingju Liu, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Philip Coleman |
ICASSP | 2 |
| 2017 | Fast tagging of natural sounds using marginal co-regularizationabstractAutomatic and fast tagging of natural sounds in audio collections is a very challenging task due to wide acoustic variations, the large number of possible tags, the incomplete and ambiguous tags provided by different labellers. To handle these problems, we use a co-regularization approach to learn a pair of classifiers on sound and text. The first classifier maps low-level audio features to a true tag list. The second classifier maps actively corrupted tags to the true tags, reducing incorrect mappings caused by low-level acoustic variations in the first classifier, and to augment the tags with additional relevant tags. Training the classifiers is implemented using marginal co-regularization, pair of which draws the two classifiers into agreement by a joint optimization. We evaluate this approach on two sound datasets, Freefield1010 and Task4 of DCASE2016. The results obtained show that marginal co-regularization outperforms the baseline GMM in both efficiency and effectiveness. Qiang Huang 0007, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 2 |
| 2017 | A joint detection-classification model for audio tagging of weakly labelled dataabstractAudio tagging aims to assign one or several tags to an audio clip. Most of the datasets are weakly labelled, which means only the tags of the clip are known, without knowing the occurrence time of the tags. The labeling of an audio clip is often based on the audio events in the clip and no event level label is provided to the user. Previous works have used the bag of frames model assume the tags occur all the time, which is not the case in practice. We propose a joint detection-classification (JDC) model to detect and classify the audio clip simultaneously. The JDC model has the ability to attend to informative and ignore uninformative sounds. Then only informative regions are used for classification. Experimental results on the “CHiME Home” dataset show that the JDC model reduces the equal error rate (EER) from 19.0% to 16.9%. More interestingly, the audio event detector is trained successfully without needing the event level label. Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley |
ICASSP | 2 |
| 2017 | Convolutional gated recurrent neural network incorporating spatial features for audio taggingabstractEnvironmental audio tagging is a newly proposed task to predict the presence or absence of a specific audio event in a chunk. Deep neural network (DNN) based methods have been successfully adopted for predicting the audio tags in the domestic audio scene. In this paper, we propose to use a convolutional neural network (CNN) to extract robust features from mel-filter banks (MFBs), spectrograms or even raw waveforms for audio tagging. Gated recurrent unit (GRU) based recurrent neural networks (RNNs) are then cascaded to model the long-term temporal structure of the audio signal. To complement the input information, an auxiliary CNN is designed to learn on the spatial features of stereo recordings. We evaluate our proposed methods on Task 4 (audio tagging) of the Detection and Classification of Acoustic Scenes and Events 2016 (DCASE 2016) challenge. Compared with our recent DNN-based method, the proposed structure can reduce the equal error rate (EER) from 0.13 to 0.11 on the development set. The spatial features can further reduce the EER to 0.10. The performance of the end-to-end learning on raw waveforms is also comparable. Finally, on the evaluation set, we get the state-of-the-art performance with 0.12 EER while the performance of the best existing system is 0.15 EER. Yong Xu 0004, Qiuqiang Kong, Qiang Huang 0007, Wenwu Wang 0001, Mark D. Plumbley |
IJCNN | 1 |
| 2017 | Attention and Localization Based on a Deep Convolutional Recurrent Model for Weakly Supervised Audio TaggingabstractAudio tagging aims to perform multi-label classification on audio \nchunks and it is a newly proposed task in the Detection and \nClassification of Acoustic Scenes and Events 2016 (DCASE \n2016) challenge. This task encourages research efforts to better \nanalyze and understand the content of the huge amounts of \naudio data on the web. The difficulty in audio tagging is that \nit only has a chunk-level label without a frame-level label. This \npaper presents a weakly supervised method to not only predict \nthe tags but also indicate the temporal locations of the occurred \nacoustic events. The attention scheme is found to be effective \nin identifying the important frames while ignoring the unrelated \nframes. The proposed framework is a deep convolutional recurrent \nmodel with two auxiliary modules: an attention module \nand a localization module. The proposed algorithm was evaluated \non the Task 4 of DCASE 2016 challenge. State-of-the-art \nperformance was achieved on the evaluation set with equal error \nrate (EER) reduced from 0.13 to 0.11, compared with the \nconvolutional recurrent baseline system. Yong Xu 0004, Qiuqiang Kong, Qiang Huang 0007, Wenwu Wang 0001, Mark D. Plumbley |
INTERSPEECH | 1 |
| 2017 | Binaural and log-power spectra features with deep neural networks for speech-noise separationabstractBinaural features of interaural level difference and interaural phase difference have proved to be very effective in training deep neural networks (DNNs), to generate time-frequency masks for target speech extraction in speech-speech mixtures. However, effectiveness of binaural features is reduced in more common speech-noise scenarios, since the noise may over-shadow the speech in adverse conditions. In addition, the reverberation also decreases the sparsity of binaural features and therefore adds difficulties to the separation task. To address the above limitations, we highlight the spectral difference between speech and noise spectra and incorporate the log-power spectra features to extend the DNN input. Tested on two different reverberant rooms at different signal to noise ratios (SNR), our proposed method shows advantages over the baseline method using only binaural features in terms of signal to distortion ratio (SDR) and Short-Time Perceptual Intelligibility (STOI). Alfredo Zermini, Qingju Liu, Yong Xu 0004, Mark D. Plumbley, Dave Betts, Wenwu Wang 0001 |
MMSP | 3 |
| 2017 | Hierarchical deep neural network for multivariate regression
Jun Du 0002, Yong Xu 0004 |
Pattern Recognit. | 2 |
| 2017 | Unsupervised Feature Learning Based on Deep Models for Environmental Audio TaggingabstractEnvironmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper, we make contributions to audio tagging in two parts, respectively, acoustic modeling and feature learning. We propose to use a shrinking deep neural network (DNN) framework incorporating unsupervised feature learning to handle the multilabel classification task. For the acoustic modeling, a large set of contextual frames of the chunk are fed into the DNN to perform a multilabel classification for the expected tags, considering that only chunk (or utterance) level rather than frame-level labels are available. Dropout and background noise aware training are also adopted to improve the generalization capability of the DNNs. For the unsupervised feature learning, we propose to use a symmetric or asymmetric deep denoising auto-encoder (syDAE or asyDAE) to generate new data-driven features from the logarithmic Mel-filter banks features. The new features, which are smoothed against background noise and more compact with contextual information, can further improve the performance of the DNN baseline. Compared with the standard Gaussian mixture model baseline of the DCASE 2016 audio tagging challenge, our proposed method obtains a significant equal error rate (EER) reduction from 0.21 to 0.13 on the development set. The proposed asyDAE system can get a relative 6.7% EER reduction compared with the strong DNN baseline on the development set. Finally, the results also show that our approach obtains the state-of-the-art performance with 0.15 EER on the evaluation set of the DCASE 2016 audio tagging task while EER of the first prize of this challenge is 0.17. Yong Xu 0004, Qiang Huang 0007, Wenwu Wang 0001, Peter Foster, Siddharth Sigtia, Philip J. B. Jackson, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | DNN-based speech bandwidth expansion and its application to adding high-frequency missing features for automatic speech recognition of narrowband speech
Kehuang Li, Zhen Huang 0001, Yong Xu 0004, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2015 | Multi-objective learning and mask-based post-processing for deep neural network based speech enhancementabstractWe propose a multi-objective framework to learn both secondary targets not directly related to the intended task of speech enhancement (SE) and the primary target of the clean log-power spectra (LPS) features to be used directly for constructing the enhanced speech signals.In deep neural network (DNN) based SE we introduce an auxiliary structure to learn secondary continuous features, such as mel-frequency cepstral coefficients (MFCCs), and categorical information, such as the ideal binary mask (IBM), and integrate it into the original DNN architecture for joint optimization of all the parameters.This joint estimation scheme imposes additional constraints not available in the direct prediction of LPS, and potentially improves the learning of the primary target.Furthermore, the learned secondary information as a byproduct can be used for other purposes, e.g., the IBM-based post-processing in this work.A series of experiments show that joint LPS and MFCC learning improves the SE performance, and IBM-based post-processing further enhances listening quality of the reconstructed speech. Yong Xu 0004, Jun Du 0002, Zhen Huang 0001, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |
| 2015 | A Regression Approach to Speech Enhancement Based on Deep Neural NetworksabstractIn contrast to the conventional minimum mean square error (MMSE)-based noise reduction techniques, we propose a supervised method to enhance speech by means of finding a mapping function between noisy and clean speech signals based on deep neural networks (DNNs). In order to be able to handle a wide range of additive noises in real-world situations, a large training set that encompasses many possible combinations of speech and noise types, is first designed. A DNN architecture is then employed as a nonlinear regression function to ensure a powerful modeling capability. Several techniques have also been proposed to improve the DNN-based speech enhancement system, including global variance equalization to alleviate the over-smoothing problem of the regression model, and the dropout and noise-aware training strategies to further improve the generalization capability of DNNs to unseen noise conditions. Experimental results demonstrate that the proposed framework can achieve significant improvements in both objective and subjective measures over the conventional MMSE based technique. It is also interesting to observe that the proposed DNN approach can well suppress highly nonstationary noise, which is tough to handle in general. Furthermore, the resulting DNN model, trained with artificial synthesized data, is also effective in dealing with noisy speech data recorded in real-world scenarios without the generation of the annoying musical artifact commonly observed in conventional enhancement methods. Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Robust speech recognition with speech enhanced deep neural networks
Jun Du 0002, Qing Wang 0008, Tian Gao 0005, Yong Xu 0004, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2014 | Dynamic noise aware training for speech enhancement based on deep neural networksabstractWe propose three algorithms to address the mismatch problem in deep neural network (DNN) based speech enhancement. First, we investigate noise aware training by incorporating noise informationin the testutterance with anideal binary maskbased dynamic noise estimation approach to improve DNN’s speech separation ability from the noisy signal. Next, a set of more than 100 noise types is adopted to enrich the generalization capabilities of the DNN to unseen and non-stationary noise conditions. Finally, the quality of the enhanced speech can further be improved by global variance equalization. Empirical results show that each of the three proposed techniques contributes to the performance improvement. Compared to the conventional logarithmic minimum mean squared error speech enhancement method, our DNN system achieves 0.32 PESQ (perceptual evaluation of speech quality) improvement across six signal-tonoise ratio levels ranging from -5dB to 20dB on a test set with unknown noise types. We also observe that the combined strategies can well suppress highly non-stationary noise better than all the competing state-of-the-art techniques we have evaluated. Index Terms: Speech enhancement, deep neural networks, noise aware training, ideal binary mask, non-stationary noise Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |
| 2014 | An Experimental Study on Speech Enhancement Based on Deep Neural NetworksabstractThis letter presents a regression-based speech enhancement framework using deep neural networks (DNNs) with a multiple-layer deep architecture. In the DNN learning process, a large training set ensures a powerful modeling capability to estimate the complicated nonlinear mapping from observed noisy speech to desired clean signals. Acoustic context was found to improve the continuity of speech to be separated from the background noises successfully without the annoying musical artifact commonly observed in conventional speech enhancement algorithms. A series of pilot experiments were conducted under multi-condition training with more than 100 hours of simulated speech data, resulting in a good generalization capability even in mismatched testing conditions. When compared with the logarithmic minimum mean square error approach, the proposed DNN-based algorithm tends to achieve significant improvements in terms of various objective quality measures. Furthermore, in a subjective preference evaluation with 10 listeners, 76.35% of the subjects were found to prefer DNN-based enhanced speech to that obtained with other conventional technique. Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001 |
IEEE Signal Process. Lett. | 1 |