EDBT 2026 Demo / reviewers in the wild / expert
Meng Yu 0003
dblp:13/2440-3
· DBLP profile ↗
68ranked-venue papers
9as first author
38since 2021 · last 2026
0000-0002-0031-9156ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 8 first-author · 32 since 2021Artificial intelligence and machine learning · 43 · 9 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveabstractExisting neural vocoders have demonstrated promising performance by leveraging Mel-spectrum as an acoustic feature for conditional audio generation. Nonetheless, they remain constrained by an inherent ``performance-cost'' dilemma that significantly hinders the development of this field. This paper revisits this foundational task from a novel degradation perspective, where Mel-spectrum is regarded as a special signal degradation process from the target spectrum. Drawing inspiration from traditional sparse signal recovery problems, we propose DegVoC, a GAN-based neural vocoder with a two-step solution procedure. First, by exploiting degradation priors, we attempt to retrieve the initial spectral structure from Mel-domain representations as an initial solution via a simple linear transformation. Based on that, we introduce a deep prior solver that accounts for the heterogeneous distribution of sub-bands in the time-frequency domain. A convolution-style attention module with a large kernel size is specially devised for efficient inter-frame and inter-band contextual modeling. With 3.89 M parameters and substantially reduced inference complexity, DegVoC achieves state-of-the-art performance across objective and subjective evaluations, outperforming existing GAN-, DDPM- and flow-matching-based baselines. Andong Li, Lingling Dai, Rilin Chen, Meng Yu 0003, Xiaodong Li 0002, Dong Yu 0001, Chengshi Zheng |
AAAI | 6 |
| 2026 | Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningabstractRecent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded substantial benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of achieving human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs through improved adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that assist the model in distinguishing between valid and flawed reasoning paths during training. Experimental results demonstrate that Audio-Thinker models outperform existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities. Chenxing Li, Wenfu Wang, Hao Zhang 0112, Hualei Wang, Meng Yu 0003, Dong Yu 0001 |
AAAI | 6 |
| 2026 | Trainable multi-channel front-ends for joint beamforming and speaker embedding extractionabstractMulti-channel speaker verification (SV), employing numerous microphones for capturing enrollment and/or test recordings, gained attention for its benefits in far-field scenarios. While some studies approach the problem by designing multi-channel embedding extractors, we focus on building and thoroughly analyzing a framework integrating beamforming pre-processing paired with single-channel embedding extraction. This strategy benefits from accommodating both multi-channel and single-channel inputs. Furthermore, it provides human-interpretable intermediate output — enhanced speech — that can be independently evaluated and related to SV performance. We first focus on the front-end, taking advantage of deep-learning source separation for direct or indirect mask estimation required by the beamformer. We alternate single-channel network architectures, subsequently extended to multi-channel ones by reference channel attention (RCA). We also analyze the impact of beamformer and network output fusion. Finally, we show improvements brought by end-to-end fine-tuning the entire architecture facilitated by our newly designed multi-channel corpus, MultiSV2, extending our previous MultiSV dataset. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký, Meng Yu 0003 |
Comput. Speech Lang. | 6 |
| 2025 | Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone ArrayabstractSpatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for efficient sound field capture, which limits their flexibility in practical scenarios. We propose a deep learning (DL)-based approach, leveraging a two-stage network architecture for encoding circular microphone array signals into second-order Ambisonics (SOA) in multi-speaker environments. In addition, we introduce: (i) a novel loss function based on spatial power maps to regularize inter-channel correlations of the Ambisonic signals, and (ii) a channel permutation technique to resolve the ambiguity of encoding vertical information using a horizontal circular array. Evaluation on simulated speech and noise datasets shows that our approach consistently outperforms traditional signal processing (SP) and DL-based methods, providing significantly better timbral and spatial quality and higher source localization accuracy. Binaural audio demos with visualizations are available at https://bridgoon97.github.io/NeuralAmbisonicEncoding/. Vinay Kothapally, Meng Yu 0003, Dong Yu 0001 |
ICASSP | 3 |
| 2025 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and SynthesisabstractIn this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released. Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001 |
ICASSP | 2 |
| 2025 | BridgeVoC: Neural Vocoder with Schrödinger BridgeabstractWhile previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restoration task, this paper proposes a time-frequency (T-F) domain-based neural vocoder with the Schrödinger Bridge, called BridgeVoC, which is the first to follow the data-to-data generation paradigm. Specifically, the mel-spectrogram can be projected into the target linear-scale domain and regarded as a degraded spectral representation with a deficient rank distribution. Based on this, the Schrödinger Bridge is leveraged to establish a connection between the degraded and target data distributions. During the inference stage, starting from the degraded representation, the target spectrum can be gradually restored rather than generated from a Gaussian noise process. Quantitative experiments on LJSpeech and LibriTTS show that BridgeVoC achieves faster inference and surpasses existing diffusion-based vocoder baselines, while also matching or exceeding non-diffusion state-of-the-art methods across evaluation metrics. Rilin Chen, Meng Yu 0003, Chengshi Zheng, Dong Yu 0001, Andong Li |
IJCAI | 4 |
| 2025 | Scaling beyond Denoising: Submitted System and Findings in URGENT Challenge 2025
Zhihang Sun, Andong Li, Rilin Chen, Meng Yu 0003, Chengshi Zheng, Yi Zhou 0014, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2025 | From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language ModelsabstractThis paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications. Zhaoxi Mu, Rilin Chen, Andong Li, Meng Yu 0003, Xinyu Yang 0001, Dong Yu 0001 |
ACM Multimedia | 4 |
| 2024 | Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 4 |
| 2024 | SMRU: Split-And-Merge Recurrent-Based UNet For Acoustic Echo Cancellation And Noise SuppressionabstractThe proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down-/upsampling layer with varying compression ratios and the dualpath structure for inter- and intra-band modeling. The model is configured from $50 \mathrm{M} / \mathrm{s}$ to $6.8 \mathrm{G} / \mathrm{s}$ in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements. Zhihang Sun, Andong Li, Rilin Chen, Hao Zhang 0112, Meng Yu 0003, Yi Zhou 0014, Dong Yu 0001 |
SLT | 5 |
| 2024 | Enhanced Acoustic Howling Suppression via Hybrid Kalman Filter and Deep Learning ModelsabstractThis paper presents a comprehensive study addressing the challenging problem of acoustic howling suppression (AHS) through the fusion of Kalman filter and deep learning techniques. We introduce two integration approaches: HybridAHS, which concatenates Kalman and neural networks (NN), and NeuralKalmanAHS, where NN modules are embedded inside the Kalman filter for signal and parameter estimation. In HybridAHS, we explore two implementation methods. One is trained offline using pre-processed signals with a light training burden, while the other employs a recursive training strategy with training signals generated adaptively. The offline model serves as an initialization for recursively training the other model. With NeuralKalmanAHS, we harness the power of NN modules to refine the reference signal and improve covariance matrices estimation in the Kalman filter, resulting in enhanced feedback suppression. Our methods capitalize on the strengths of traditional and deep learning-based AHS techniques. We have explored different variants of combining Kalman filter and NN and systematically compared their howling suppression performance, providing users with versatile solutions for addressing AHS. Furthermore, by employing the proposed recursive training, we effectively mitigate the mismatch issues that plagued previous NN-based AHS methods. Extensive experimental results show the superiority of our approach over baseline techniques. Hao Zhang 0112, Yixuan Zhang 0005, Meng Yu 0003, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Neuralecho: Hybrid of Full-Band and Sub-Band Recurrent Neural Network For Acoustic Echo Cancellation and Speech EnhancementabstractThis paper presents a hybrid of full-band and sub-band recurrent neural network (RNN) model, named NeuralEcho, to jointly solve echo and noise suppression. The full-band model part processes the signal’s entire frequency bands as a whole, while the sub-band model part divides the features into sub-bands and processes each sub-band separately. This approach allows the model to capture both the fine-grained local details of the sub-band processing and the global context of the full-band processing. The single-channel model is then generalized to accommodate a range of input channel numbers. Experimental results show that the hybrid model outperforms the conventional full-band models in terms of objective speech quality metrics and speech recognition accuracy. This suggests that the hybrid approach of full-band and sub-band processing can be a promising direction for future research in the field of speech enhancement. Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
ASRU | 1 |
| 2023 | Neuralkalman: A Learnable Kalman Filter for Acoustic Echo CancellationabstractThe robustness of the Kalman filter to double talk and its rapid convergence make it a popular approach for addressing acoustic echo cancellation (AEC) challenges. However, the inability to model nonlinearity and the need to tune control parameters cast limitations on such adaptive filtering algorithms. In this paper, we integrate the frequency domain Kalman filter (FDKF) and deep neural networks (DNNs) into a hybrid method, called NeuralKalman, to leverage the advantages of deep learning and adaptive filtering algorithms. Specifically, we employ a DNN to estimate nonlinearly distorted far-end signals, a transition factor, and the nonlinear transition function in the state equation of the FDKF algorithm. Experimental results show that the proposed NeuralKalman improves the performance of FDKF significantly and outperforms strong baseline methods. Yixuan Zhang 0005, Meng Yu 0003, Hao Zhang 0112, Dong Yu 0001, DeLiang Wang |
ASRU | 2 |
| 2023 | Deep Neural Mel-Subband Beamformer for in-Car Speech SeparationabstractWhile current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies independently which results in higher computational costs and inference times, making them unsuitable for real-world use. In this paper, we propose DL-based mel-subband spatio-temporal beamformer to perform speech separation in a car environment with reduced computation cost and inference time. As opposed to conventional subband (SB) approaches, our framework uses a mel-scale based subband selection strategy which ensures a fine-grained processing for lower frequencies where most speech formant structure is present, and coarse-grained processing for higher frequencies. In a recursive way, robust frame-level beamforming weights are determined for each speaker location/zone in a car from the estimated subband speech and noise covariance matrices. Furthermore, proposed framework also estimates and suppresses any echoes from the loudspeaker(s) by using the echo reference signals. We compare the performance of our proposed framework to several NB, SB, and full-band (FB) processing techniques in terms of speech quality and recognition metrics. Based on experimental evaluations on simulated and real-world recordings, we find that our proposed framework achieves better separation performance over all SB and FB approaches and achieves performance closer to NB processing techniques while requiring lower computing cost. Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
ICASSP | 3 |
| 2023 | Zoneformer: On-device Neural Beamformer For In-car Multi-zone Speech Separation, Enhancement and Echo Cancellation
Yong Xu 0004, Vinay Kothapally, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
INTERSPEECH | 3 |
| 2023 | Hybrid AHS: A Hybrid of Kalman Filter and Deep Learning for Acoustic Howling Suppression
Hao Zhang 0112, Meng Yu 0003, Yuzhong Wu, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2022 | Fast-Rir: Fast Neural Diffuse Room Impulse Response GeneratorabstractWe present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RIR takes rectangular room dimensions, listener and speaker positions, and reverberation time (T60) as inputs and generates specular and diffuse reflections for a given acoustic environment. Our FAST-RIR is capable of generating RIRs for a given input T60with an average error of 0.02s. We evaluate our generated RIRs in automatic speech recognition (ASR) applications using Google Speech API, Microsoft Speech API, and Kaldi tools. We show that our proposed FAST-RIR with batch size 1 is 400 times faster than a state-of-the-art diffuse acoustic simulator (DAS) on a CPU and gives similar performance to DAS in ASR experiments. Our FAST-RIR is 12 times faster than an existing GPU-based RIR generator (gpuRIR). We show that our FAST-RIR outperforms gpuRIR by 2.5% in an AMI far-field ASR benchmark. Anton Ratnarajah, Shixiong Zhang 0001, Meng Yu 0003, Zhenyu Tang 0001, Dinesh Manocha, Dong Yu 0001 |
ICASSP | 3 |
| 2022 | Joint Modeling of Code-Switched and Monolingual ASR via Conditional FactorizationabstractConversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a general framework to jointly model the likelihoods of the monolingual and code-switch sub-tasks that comprise bilingual speech recognition. By defining the monolingual sub-tasks with label-to-frame synchronization, our joint modeling framework can be conditionally factorized such that the final bilingual output, which may or may not be code-switched, is obtained given only monolingual information. We show that this conditionally factorized joint framework can be modeled by an end-to-end differentiable neural network. We demonstrate the efficacy of our proposed model on bilingual Mandarin-English speech recognition across both monolingual and code-switched corpora. Brian Yan, Meng Yu 0003, Shixiong Zhang 0001, Siddharth Dalmia, Dan Berrebbi, Chao Weng, Shinji Watanabe 0001, Dong Yu 0001 |
ICASSP | 3 |
| 2022 | Towards end-to-end Speaker Diarization with Generalized Neural Speaker ClusteringabstractSpeaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarization systems are complicated and sometimes lack of satisfying generalization capabilities. In this study, we present a novel speaker diarization system, with a generalized neural speaker clustering module as the backbone. The whole system can be simplified to contain only two major parts, a speaker embedding extractor followed by a clustering module. Both parts are implemented with neural networks. In the training phase, an on-the-fly spoken dialogue generator is designed to provide the system with audio streams and the corresponding annotations in categories of non-speech, overlapped speech and active speakers. The chunk-wise inference and a speaker verification based tracing module are conducted to handle the arbitrary number of speakers. We demonstrate that the proposed speaker diarization system is able to integrate SAD, OSD and speaker segmentation/clustering, and yield competitive results in the VoxConverse20 benchmarks. Jiatong Shi, Chao Weng, Meng Yu 0003, Dong Yu 0001 |
ICASSP | 4 |
| 2022 | Joint Neural AEC and Beamforming with Double-Talk DetectionabstractAcoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback.However, nonlinear distortions induced by audio devices, background noise, reverberation, and double-talk reduce the efficiency of conventional AEC systems.Several hybrid AEC models were proposed to address this, which use deep learning models to suppress residual echo from standard adaptive filtering.This paper proposes deep learning-based joint AEC and beamforming model (JAECBF) building on our previous self-attentive recurrent neural network (RNN) beamformer.The proposed network consists of two modules: (i) multi-channel neural-AEC, and (ii) joint AEC-RNN beamformer with a double-talk detection (DTD) that computes time-frequency (T-F) beamforming weights.We train the proposed model in an end-to-end approach to eliminate background noise and echoes from far-end audio devices, which include nonlinear distortions.From experimental evaluations, we find the proposed network outperforms other multi-channel AEC and denoising systems in terms of speech recognition rate and overall speech quality. Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
INTERSPEECH | 3 |
| 2022 | EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of SpeakersabstractIn this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neural diarization (EEND) models, speaker counting with encoder-decoder based attractors (EDA), and speech separation using Conv-TasNet. In addition, we propose a multiple$1 \times 1$convolutional layer architecture for estimating the separation masks corresponding to a flexible number of speakers and a fusion technique for refining the separated speech signal with obtained speaker diarization information to improve the joint framework. Experiments using the LibriMix dataset show that our proposed method outperforms the single-task baselines in both diarization and separation metrics for fixed and flexible numbers of speakers and improves speaker counting performance for flexible numbers of speakers. All materials will be open-sourced and reproducible in ESPnet toolkit11https://github.com/espnet/espnet. Soumi Maiti, Yushi Ueda, Shinji Watanabe 0001, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004 |
SLT | 5 |
| 2022 | An investigation of neural uncertainty estimation for target speaker extraction equipped RNN transducer
Jiatong Shi, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001 |
Comput. Speech Lang. | 5 |
| 2022 | Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognitionabstractMulti-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural networks to estimate the direction of arrival (DOA) of all the speakers simultaneously from the audio mixture. At the heart of the proposal is a source splitting mechanism that creates source-specific intermediate representations inside the network. This allows our model to give source-specific posteriors as the output unlike the traditional multi-label classification approach . Existing deep learning methods perform a frame level prediction, whereas our approach performs an utterance level prediction by incorporating temporal selection and averaging inside the network to avoid post-processing. We also experiment with various loss functions and show that a variant of earth mover distance (EMD) is very effective in classifying DOA at a very high resolution by modeling inter-class relationships. In addition to using the prediction error as a metric for evaluating our localization model, we also establish its potency as a frontend with automatic speech recognition (ASR) as the downstream task. We convert the estimated DOAs into a feature suitable for ASR and pass it as an additional input feature to a strong multi-channel and multi-talker speech recognition baseline. This added input feature drastically improves the ASR performance and gives a word error rate (WER) of 6.3% on the evaluation data of our simulated noisy two speaker mixtures, while the baseline which does not use explicit localization input has a WER of 11.5%. We also perform ASR evaluation on real recordings with the overlapped set of the MC-WSJ-AV corpus in addition to simulated mixtures. Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001 |
Comput. Speech Lang. | 4 |
| 2021 | 3D Spatial Features for Multi-Channel Target Speech SeparationabstractThe use of speaker's directional information for speech sepa-ration and speech recognition has demonstrated the state-of-the-art performances on multi-talker scenarios. One major limitation of previous approaches using speaker's directional information is the significant performance degradation when the coming directions of two sound sources are close. To address these challenges, this paper proposed a set of new three-dimensional (3D) spatial features for target speech sep-aration, by leveraging all the 3D location information of the target speaker, including azimuth, elevation, and the distance to the microphone array center. Previous works in this area are extended in two important directions. First, the traditional 1D directional features are generalized to 3D spatial features. Thus more discriminative spatial diversity between speakers is achieved. Second, to unleash the full power of these 3D spatial features, a microphone pair-wise attention model is also proposed. The proposed features and models were evaluated on both simulated reverberant datasets and real recordings under near and far-field conditions. Exper-imental results show that both proposed 3D spatial features and attention models can significantly improve the separation performance as well as reducing the recognition error rate. Rongzhi Gu, Shixiong Zhang 0001, Meng Yu 0003, Dong Yu 0001 |
ASRU | 3 |
| 2021 | Improving RNN Transducer with Target Speaker Extraction and Neural Uncertainty EstimationabstractTarget-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech extraction and Recurrent Neural Network Transducer (RNN-T). To stabilize the joint-training, we propose a multi-stage training strategy that pre-trains and fine-tunes each module in the system before joint-training. Meanwhile, speaker identity and speech enhancement uncertainty measures are proposed to compensate for residual noise and artifacts from the target speech extraction module. Compared to a recognizer fine-tuned with a target speech extraction model, our experiments show that adding the neural uncertainty module significantly reduces 17% relative Character Error Rate (CER) on multi-speaker signals with background noise. The multi-condition experiments indicate that our method can achieve 9% relative performance gain in the noisy condition while maintaining the performance in the clean condition. Jiatong Shi, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001 |
ICASSP | 5 |
| 2021 | Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source LocalizationabstractThis paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end (E2E) neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the azimuth angle of the sources with respect to the microphone array is defined as a latent variable. This angle controls the quality of separation, which in turn determines the ASR performance. All three functionalities of D-ASR: localization, separation, and recognition are connected as a single differentiable neural network and trained solely based on ASR error minimization objectives. The advantages of D-ASR over existing methods are threefold: (1) it provides explicit speaker locations, (2) it improves the explainability factor, and (3) it achieves better ASR performance as the process is more streamlined. In addition, D-ASR does not require explicit direction of arrival (DOA) supervision like existing data-driven localization models, which makes it more appropriate for realistic data. For the case of two source mixtures, D-ASR achieves an average DOA prediction error of less than three degrees. It also outperforms a strong far-field multi-speaker end-to-end system in both separation quality and ASR performance. Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
ICASSP | 4 |
| 2021 | Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive LearningabstractIn this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a momentum contrastive (MoCo) learning framework, where the MoCo speaker embedding system utilizes a queue to maintain a large set of negative examples. We show that better speaker embeddings can be learned by momentum contrastive learning. Next, alternative augmentation strategies are explored to normalize extrinsic speaker variabilities of two random segments from the same speech utterance. Specifically, augmentation in the waveform largely improves the speaker representations for SV tasks. The proposed MoCo speaker embedding is further improved when a prototypical memory bank is introduced, which encourages the speaker embeddings to be closer to their assigned prototypes with an intermediate clustering step. In addition, we generalize the self-supervised framework to a semi-supervised scenario where only a small portion of the data is labeled. Comprehensive experiments on the Voxceleb dataset demonstrate that our proposed self-supervised approach achieves competitive performance compared with existing techniques, and can approach fully supervised results with partially labeled data. Chao Weng, Meng Yu 0003, Dong Yu 0001 |
ICASSP | 4 |
| 2021 | ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech SeparationabstractSpeech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech recognition (ASR) systems. The conventional mask-based minimum variance distortionless response (MVDR) beamformer can be used to minimize the distortion, but comes with high level of residual noise. Furthermore, the matrix operations (e.g., matrix inversion) involved in the conventional MVDR solution are sometimes numerically unstable when jointly trained with neural networks. In this paper, we propose a novel all deep learning MVDR framework, where the matrix inversion and eigenvalue decomposition are replaced by two recurrent neural networks (RNNs), to resolve both issues at the same time. The proposed method can greatly reduce the residual noise while keeping the target speech undistorted by leveraging on the RNN-predicted frame-wise beamforming weights. The system is evaluated on a Mandarin audio-visual corpus and compared against several state-of-the-art (SOTA) speech separation systems. Experimental results demonstrate the superiority of the proposed method across several objective metrics and ASR accuracy. Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001 |
ICASSP | 3 |
| 2021 | Towards Robust Speaker Verification with Target Speaker EnhancementabstractThis paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV). Specifically, an enrollment speaker conditioned speech enhancement module is employed as the front-end for extracting target speaker from its mixture with interfering speakers and environmental noises. Compared with the conventional target speaker enhancement models, nontarget speaker/interference suppression should draw additional attention for SV. Therefore, an effective nontarget speaker sampling strategy is explored. To improve speaker embedding extraction with a light-weighted model, a teacher-student (T/S) training is proposed to distill speaker discriminative information from large models to small models. Iterative inference is investigated to address the noisy speaker enrollment problem. We evaluate the proposed method on two SV tasks, i.e., one heavily overlapped speech and the other one with comprehensive noise types in vehicle environments. Experiments show significant and consistent improvements in Equal Error Rate (EER) over the state-of-the-art baselines. Meng Yu 0003, Chao Weng, Dong Yu 0001 |
ICASSP | 2 |
| 2021 | A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped SpeechabstractIn multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint training framework of the front-end multi-look speech separator and the back-end speaker embedding extractor is proposed for multi-channel overlapped speech. To better leverage the complementarity between the speech separator and the speaker embedding extractor, several training strategies are proposed to jointly optimize the two modules. Experimental results show that the proposed joint training framework significantly outperforms the individual SV system by around 52% relative EER reduction. Additionally, the robustness of the proposed framework is further evaluated under different conditions. Naijun Zheng, Na Li 0012, Bo Wu 0011, Meng Yu 0003, Jianwei Yu 0001, Chao Weng, Dan Su 0002, Xunying Liu, Helen M. Meng |
ICASSP | 4 |
| 2021 | MIMO Self-Attentive RNN Beamformer for Multi-Speaker Speech SeparationabstractRecently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion and eigenvalue decomposition with two RNNs.In this work, we present a self-attentive RNN beamformer to further improve our previous RNN-based beamformer by leveraging on the powerful modeling capability of self-attention.Temporal-spatial self-attention module is proposed to better learn the beamforming weights from the speech and noise spatial covariance matrices.The temporal self-attention module could help RNN to learn global statistics of covariance matrices.The spatial self-attention module is designed to attend on the cross-channel correlation in the covariance matrices.Furthermore, a multi-channel input with multi-speaker directional features and multi-speaker speech separation outputs (MIMO) model is developed to improve the inference efficiency.The evaluations demonstrate that our proposed MIMO self-attentive RNN beamformer improves both the automatic speech recognition (ASR) accuracy and the perceptual estimation of speech quality (PESQ) against prior arts. Xiyun Li, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Jiaming Xu 0001, Bo Xu 0002, Dong Yu 0001 |
Interspeech | 3 |
| 2021 | TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech DereverberationabstractIn this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach on the deep neural network (DNN) for environment-aware speech dereverberation, which can adaptively attend to the contextual information. More specifically, a FullBand based Temporal Attention approach (FTA) is proposed, which models the correlations between the fullband information of the context frames. In addition, considering the difference between the attenuation of high frequency bands and low frequency bands (high frequency bands attenuate faster than low frequency bands) in the room impulse response (RIR), we also propose a SubBand based Temporal Attention approach (STA). In order to guide the network to be more aware of the reverberant environments, we jointly optimize the dereverberation network and the reverberation time (RT60) estimator in a multi-task manner. Our experimental results indicate that the proposed method outperforms our previously proposed reverberation-time-aware DNN and the learned attention weights are fully physical consistent. We also report a preliminary yet promising dereverberation and recognition experiment on real test data. Helin Wang, Bo Wu 0011, Lianwu Chen, Meng Yu 0003, Jianwei Yu 0001, Yong Xu 0004, Shixiong Zhang 0001, Chao Weng, Dan Su 0002, Dong Yu 0001 |
Interspeech | 4 |
| 2021 | Generalized Spatio-Temporal RNN Beamformer for Target Speech SeparationabstractAlthough the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal recurrent neural network based beamformer (RNN-BF) for target speech separation. This new beamforming framework directly learns the beamforming weights from the estimated speech and noise spatial covariance matrices. Leveraging on the temporal modeling capability of RNNs, the RNN-BF could automatically accumulate the statistics of the speech and noise covariance matrices to learn the frame-level beamforming weights in a recursive way. An RNN-based generalized eigenvalue (RNN-GEV) beamformer and a more generalized RNN beamformer (GRNN-BF) are proposed. We further improve the RNN-GEV and the GRNN-BF by using layer normalization to replace the commonly used mask normalization on the covariance matrices. The proposed GRNN-BF obtains better performance against prior arts in terms of speech quality (PESQ), speech-to-noise ratio (SNR) and word error rate (WER). Yong Xu 0004, Zhuohuang Zhang, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001 |
Interspeech | 3 |
| 2021 | MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality AssessmentabstractThe objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion score (MOS) tests. Non-intrusive speech quality assessment has attracted much attention recently due to the lack of access to clean reference signals for objective evaluations in real scenarios. In this paper, we propose a novel non-intrusive speech quality measurement model, MetricNet, which leverages label distribution learning and joint speech reconstruction learning to achieve significantly improved performance compared to the existing non-intrusive speech quality measurement models. We demonstrate that the proposed approach yields promisingly high correlation to the intrusive objective evaluation of speech quality on clean, noisy and processed speech data. Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
Interspeech | 1 |
| 2021 | Neural Mask based Multi-channel Convolutional Beamforming for Joint Dereverberation, Echo Cancellation and DenoisingabstractThis paper proposes a new joint optimization framework for simultaneous dereverberation, acoustic echo cancellation, and denoising, which is motivated by the recently proposed con-volutional beamformer for simultaneous denoising and dereverberation. Using the echo aware mask based beamforming framework, the proposed algorithm could effectively deal with double-talk case and local inference, etc. The evaluations based on ERLE for echo only, and PESQ for double-talk demonstrate that the proposed algorithm could significantly improve the performance. Meng Yu 0003, Yong Xu 0004, Chao Weng, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001 |
SLT | 2 |
| 2021 | WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and DereverberationabstractThis paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted Power minimization Distortionless response (WPD) beamformer can perform separation and dereverberation simultaneously, the noise cancellation component still has the potential to progress. We propose an improved neural WPD beamformer called "WPD++" by an enhanced beamforming module in the conventional WPD and a multi-objective loss function for the joint training. The beamforming module is improved by utilizing the spatio-temporal correlation. A multi-objective loss, including the complex spectra domain scale-invariant signal-to-noise ratio (C-Si-SNR) and the magnitude domain mean square error (Mag-MSE), is properly designed to make multiple constraints on the enhanced speech and the desired power of the dry clean signal. Joint training is conducted to optimize the complex-valued mask estimator and the WPD++ beamformer in an end-to-end way. The results show that the proposed WPD++ outperforms several state-of-the-art beamformers on the enhanced speech quality and word error rate (WER) of ASR. Zhaoheng Ni, Yong Xu 0004, Meng Yu 0003, Bo Wu 0011, Shixiong Zhang 0001, Dong Yu 0001, Michael I. Mandel |
SLT | 3 |
| 2021 | An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and SeparationabstractSpeech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance. Daniel Michelsanti, Zheng-Hua Tan, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Jesper Jensen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Multi-Channel Multi-Frame ADL-MVDR for Target Speech SeparationabstractMany purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems. Minimum variance distortionless response (MVDR) filters are often adopted to remove nonlinear distortions, however, conventional neural mask-based MVDR systems still result in relatively high levels of residual noise. Moreover, the matrix inverse involved in the MVDR solution is sometimes numerically unstable during joint training with neural networks. In this study, we propose a multi-channel multi-frame (MCMF) all deep learning (ADL)-MVDR approach for target speech separation, which extends our preliminary multi-channel ADL-MVDR approach. The proposed MCMF ADL-MVDR system addresses linear and nonlinear distortions. Spatio-temporal cross correlations are also fully utilized in the proposed approach. The proposed systems are evaluated using a Mandarin audio-visual corpus and are compared with several state-of-the-art approaches. Experimental results demonstrate the superiority of our proposed systems under different scenarios and across several objective evaluation metrics, including ASR performance. Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Donald S. Williamson, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature LearningabstractHand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In this work, we propose an integrated architecture for learning spatial features directly from the multi-channel speech waveforms within an end-to-end speech separation framework. In this architecture, time-domain filters spanning signal channels are trained to perform adaptive spatial filtering. These filters are implemented by a 2d convolution (conv2d) layer and their parameters are optimized using a speech separation objective function in a purely data-driven fashion. Furthermore, inspired by the IPD formulation, we design a conv2d kernel to compute the inter-channel convolution differences (ICDs), which are expected to provide the spatial cues that help to distinguish the directional sources. Evaluation results on simulated multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed ICD based MCSS model improves the overall signal-to-distortion ratio by 10.4% over the IPD based MCSS model. Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001 |
ICASSP | 5 |
| 2020 | Integration of Multi-Look Beamformers for Multi-Channel Keyword SpottingabstractKeyword spotting (KWS) is in great demand in smart devices in the era of Internet of Things. Albeit recent progresses, the performance of KWS, measured in false alarms and false rejects, may still degrade significantly under the far field and noisy conditions. In this paper, we propose integrating multiple beamformed signals and a microphone signal as input to an end-to-end KWS model and leveraging the attention mechanism to dynamically tune the model’s attention to the reliable input sources. We demonstrate, on our large simulated and recorded noisy and far-field evaluation sets, that our proposed approach significantly improves the KWS performance and reduces the computation cost against the baseline KWS systems. Meng Yu 0003, Jie Chen 0057, Jimeng Zheng, Dan Su 0002, Dong Yu 0001 |
ICASSP | 2 |
| 2020 | Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding ExtractionabstractDeep learning based speech separation approaches have received great interest, among which the recent speaker-aware speech enhancement methods are promising for solving difficulties such as arbitrary source permutation and unknown number of sources. In this paper, we propose a novel training framework which jointly learns the speaker-conditioned target speaker extraction model and its associated speaker embedding model. The resulting unified model directly learns the appropriate speaker embedding for improved target speech enhancement. We demonstrate, on our large simulated noisy and far-field evaluation sets of overlapped speech signals, that our proposed approach significantly improves the speech enhancement performance compared to the baseline speaker-aware speech enhancement models. Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 2 |
| 2020 | Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition ObjectivesabstractTarget speech extraction is a specific case of source separation where an auxiliary information like the location or some pre-saved anchor speech examples of the target speaker is used to resolve the permutation ambiguity. Traditionally such systems are optimized based on signal reconstruction objectives. Recently end-to-end automatic speech recognition (ASR) methods have enabled to optimize source separation systems with only the transcription based objective. This paper proposes a method to jointly optimize a location guided target speech extraction module along with a speech recognition module only with ASR error minimization criteria. Experimental comparisons with corresponding conventional pipeline systems verify that this task can be realized by end-to-end ASR training objectives without using parallel clean data. We show promising target speech recognition results in mixtures of two speakers and noise, and discuss interesting properties of the proposed system in terms of speech enhancement/separation objectives and word error rates. Finally, we design a system that can take both location and anchor speech as input at the same time and show that the performance can be further improved. Aswin Shanmugam Subramanian, Chao Weng, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004, Shinji Watanabe 0001, Dong Yu 0001 |
ICASSP | 3 |
| 2020 | End-to-End Multi-Look Keyword SpottingabstractThe performance of keyword spotting (KWS), measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. In this paper, we propose a multi-look neural network modeling for speech enhancement which simultaneously steers to listen to multiple sampled look directions. The multi-look enhancement is then jointly trained with KWS to form an end-to-end KWS model which integrates the enhanced signals from multiple look directions and leverages an attention mechanism to dynamically tune the model's attention to the reliable sources. We demonstrate, on our large noisy and far-field evaluation sets, that the proposed approach significantly improves the KWS performance against the baseline KWS system and a recent beamformer based multi-beam KWS system. Meng Yu 0003, Bo Wu 0011, Dan Su 0002, Dong Yu 0001 |
INTERSPEECH | 1 |
| 2020 | Neural Spatio-Temporal Beamformer for Target Speech SeparationabstractPurely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR).On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability.In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement.Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the intermicrophone correlation that is already utilized in prior arts.Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN.The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts. Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Chao Weng, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2020 | DurIAN: Duration Informed Attention Network for Speech Synthesis
Chengzhu Yu, Heng Lu 0004, Na Hu, Meng Yu 0003, Chao Weng, Kun Xu 0005, Deyi Tuo, Shiyin Kang, Guangzhi Lei, Dan Su 0002, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2020 | Audio-Visual Multi-Channel Recognition of Overlapped SpeechabstractAutomatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption, this paper presents an audio-visual multi-channel overlapped speech recognition system featuring tightly integrated separation front-end and recognition back-end. A series of audio-visual multi-channel speech separation front-end components based on \textit{TF masking}, \textit{filter\&sum} and \textit{mask-based MVDR} beamforming approaches were developed. To reduce the error cost mismatch between the separation and recognition components, they were jointly fine-tuned using the connectionist temporal classification (CTC) loss function, or a multi-task criterion interpolation with scale-invariant signal to noise ratio (Si-SNR) error cost. Experiments suggest that the proposed multi-channel AVSR system outperforms the baseline audio-only ASR system by up to 6.81\% (26.83\% relative) and 22.22\% (56.87\% relative) absolute word error rate (WER) reduction on overlapped speech constructed using either simulation or replaying of the lipreading sentence 2 (LRS2) dataset respectively. Jianwei Yu 0001, Bo Wu 0011, Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Dong Yu 0001, Xunying Liu, Helen M. Meng |
INTERSPEECH | 7 |
| 2019 | Syllable-Dependent Discriminative Learning for Small Footprint Text-Dependent Speaker VerificationabstractThis study proposes a novel scheme of syllable-dependent discriminative speaker embedding learning for small footprint text-dependent speaker verification systems. To suppress undesired syllable variation and enhance the power of discrimination inherited in the frame-level features, we design a novel syllable-dependent clustering loss to optimize the network. Specifically, this loss function utilizes syllable labels as auxiliary supervision information to explicitly maximize inter-syllable divisibility and intra-syllable compactness between the learned frame-level features. Successively, we propose two syllable-dependent pooling mechanisms to aggregate the frame-level features to several syllable-level features by averaging those features corresponding to each syllable. The utterance-level speaker embeddings with powerful discrimination are then obtained by concatenating the syllable-level features. Experimental results on Tencent voice wake-up dataset show that our proposed scheme can accelerate the network convergence and achieve significant performance improvement against the state-of-the-art methods. Junyi Peng, Yuexian Zou, Na Li 0012, Deyi Tuo, Dan Su 0002, Meng Yu 0003, Dong Yu 0001 |
ASRU | 6 |
| 2019 | Time Domain Audio Visual Speech SeparationabstractAudio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker extraction from monaural mixtures. The architecture generalizes the previous TasNet (time-domain speech separation network) to enable multi-modal learning and at meanwhile it extends the classical audio-visual speech separation from frequency-domain to time-domain. The main components of proposed architecture include an audio encoder, a video encoder that extracts lip embedding from video streams, a multi-modal separation network and an audio decoder. Experiments on simulated mixtures based on recently released LRS2 dataset show that our method can bring 3dB+ and 4dB+ Si-SNR improvements on two- and three-speaker cases respectively, compared to audio-only TasNet and frequency-domain audio-visual networks. Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001 |
ASRU | 5 |
| 2019 | Improving Speech Enhancement with Phonetic Embedding FeaturesabstractIn this paper, we present a speech enhancement framework that leverages phonetic information obtained from the acoustic model. It consists of two separate components: (i) a long short-term memory recurrent neural network (LSTM-RNN) based speech enhancement model that takes the combination of log-power spectra (LPS) and phonetic embedding features as input to predict the complex ideal ratio mask (cIRM); and (ii) a convolutional, long short-term memory and fully connected deep neural network (CLDNN) based acoustic model that extracts the phonetic feature vector in the hidden units of its LSTM layer. Our experimental results show that the proposed framework outperforms both the conventional and phoneme-dependent speech enhancement systems under various noisy conditions, generalizes well to unseen conditions, and performs robustly to the speech interference. We further demonstrate its superior enhancement performance on unvoiced speech and report a preliminary yet promising recognition experiment on real test data. Bo Wu 0011, Meng Yu 0003, Lianwu Chen, Mingjie Jin, Dan Su 0002, Dong Yu 0001 |
ASRU | 2 |
| 2019 | Multi-band PIT and Model Integration for Improved Multi-channel Speech SeparationabstractThe recent exploration of deep learning for supervised speech separation has significantly accelerated the progress on the multi-talker speech separation problem. Multi-channel extension has attracted much research attention due to the benefit of spatial information in far-field acoustic environments. In this paper, We review the most recent models of multi-channel permutation invariant training (PIT), investigate spatial features formed by microphone pairs and their underlying impact and issue, present a multi-band architecture for effective feature encoding, and conduct a model integration between single-channel and multi-channel PIT for resolving the spatial overlapping problem in the conventional multi-channel PIT framework. The evaluation confirms the significant improvement achieved with the proposed model and training approach for the multi-channel speech separation. Lianwu Chen, Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 2 |
| 2019 | Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker VerificationabstractDeep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large Margin Cosine Loss (LMCL) in face recognition, we propose an enhanced LMCL named boundary discriminative LMCL (BD-LMCL) to emphasize the discriminative information inherited in the speaker boundaries. Unlike LMCL, where all training samples contribute equally for the objective function, only the samples around the speaker boundaries are considered during the network training with BD-LMCL. Specifically, those samples close to the boundaries are dynamically selected using top-k zero-one loss. Experimental results on a short duration corpus Android Cellphone and NIST SRE 2012 demonstrate better performance compared to LMCL and other popular loss functions. Rongjin Li, Na Li 0012, Deyi Tuo, Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 4 |
| 2019 | Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust AsrabstractIn this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covariance matrix for the beamformer. Minimum Variance Distortionless Response (MVDR) beamformer and Generalized Eigenvalue (GEV) beamformer are both investigated under the CRM-based joint training architecture. We also propose a robust mask pooling strategy among multiple channels. A long short-term memory (LSTM) based language model is utilized to re-score hypotheses which further improves the overall performance. We evaluate the proposed methods on CHiME-4 challenge dataset. The CRM based system achieves a relative 10% reduction on word error rate (WER) compared with the IRM based system. Without sequence discriminative training, our best single system already achieves an average WER 2.72% on the test set which is comparable to the state-of-the-art. Yong Xu 0004, Chao Weng, Like Hui, Meng Yu 0003, Dan Su 0002, Dong Yu 0001 |
ICASSP | 5 |
| 2019 | Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker VerificationabstractIn this paper, we present a Sequence-to-Sequence Attentional Siamese Neural Network (Seq2Seq-ASNN) that leverages temporal alignment information for end-to-end speaker verification. In prior works of speaker discriminative neural networks, utterance-level evaluation/enrollment speaker representations are usually calculated. Our proposed model, utilizing a sequence-to-sequence (Seq2Seq) attention mechanism, maps the frame-level evaluation representation into enrollment feature domain and further generates an utterance-level evaluation-enrollment joint vector for final similarity measure. Feature learning, attention mechanism, and metric learning are jointly optimized using an end-to-end loss function. Experimental results show that our proposed model outperforms various baseline methods, including the traditional i-Vector/PLDA method, multi-enrollment end-to-end speaker verification models, d-vector approaches, and a self attention model, for text-dependent speaker verification on a Tencent internal voice wake-up dataset. Meng Yu 0003, Na Li 0012, Chengzhu Yu, Jia Cui, Dong Yu 0001 |
ICASSP | 2 |
| 2019 | A Comprehensive Study of Speech Separation: Spectrogram vs Waveform SeparationabstractSpeech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain.Recently, a raw audio waveform separation network (TasNet) is introduced for single-channel data, with achieving high Si-SNR (scale-invariant source-to-noise ratio) and SDR (sourceto-distortion ratio) comparing against the state-of-the-art solution in frequency-domain.In this study, we incorporate effective components of the TasNet into a frequency-domain separation method.We compare both for alternative scenarios.We introduce a solution for directly optimizing the separation criterion in frequency-domain networks.In addition to speech separation objective and subjective measurements, we evaluate the separation performance on a speech recognition task as well.We study the speech separation problem for far-field data (more similar to naturalistic audio streams) and develop multi-channel solutions for both frequency and time-domain separators with utilizing spectral, spatial and speaker location information.For our experiments, we simulated multi-channel spatialized reverberate WSJ0-2mix dataset.Our experimental results show that spectrogram separation can achieve competitive performance with better network design.Multi-channel framework as well is shown to improve the single-channel performance relatively up to +35.5% and +46% in terms of WER and SDR, respectively. Fahimeh Bahmaninezhad, Jian Wu 0027, Rongzhi Gu, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001 |
INTERSPEECH | 6 |
| 2019 | Neural Spatial Filter: Target Speaker Speech Separation Assisted with Directional Information
Rongzhi Gu, Lianwu Chen, Shixiong Zhang 0001, Jimeng Zheng, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001 |
INTERSPEECH | 6 |
| 2019 | Direction-Aware Speaker Beam for Multi-Channel Speaker Extraction
Guanjun Li, Shan Liang 0007, Shuai Nie 0001, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li |
INTERSPEECH | 5 |
| 2019 | Jointly Adversarial Enhancement Training for Robust End-to-End Speech Recognition
Bin Liu 0041, Shuai Nie 0001, Shan Liang 0007, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li |
INTERSPEECH | 5 |
| 2019 | Improved Speaker-Dependent Separation for CHiME-5 ChallengeabstractThis paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-channel, highly-overlapped conversational speech recognition in a dinner party scenario with reverberations and nonstationary noises.We adopt a speaker-aware training method by using i-vector as the target speaker information for multi-talker speech separation.With only one unified separation model for all speakers, we achieve a 10% absolute improvement in terms of word error rate (WER) over the previous baseline of 80.28% on the development set by leveraging our newly proposed data processing techniques and beamforming approach.With our improved back-end acoustic model, we further reduce WER to 60.15% which surpasses the result of our submitted CHiME-5 challenge system without applying any fusion techniques. Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2018 | Permutation Invariant Training of Generative Adversarial Network for Monaural Speech Separation
Lianwu Chen, Meng Yu 0003, Yanmin Qian, Dan Su 0002, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2018 | Deep Extractor Network for Target Speaker Recovery from Single Channel Speech MixturesabstractSpeaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources.However, it remains challenging to achieve satisfying performance provided a very short available target speaker utterance (anchor).Here we present a novel "deep extractor network" which creates an extractor point for the target speaker in a canonical high dimensional embedding space, and pulls together the time-frequency bins corresponding to the target speaker.The proposed model is different from prior works in that the canonical embedding space encodes knowledges of both the anchor and the mixture during an end-to-end training phase: First, embeddings for the anchor and mixture speech are separately constructed in a primary embedding space, and then combined as an input to feed-forward layers to transform to a canonical embedding space which we discover more stable than the primary one.Experimental results show that given a very short utterance, the proposed model can efficiently recover high quality target speech from a mixture, which outperforms various baseline models, with 5.2% and 6.6% relative improvements in SDR and PESQ respectively compared with a baseline oracle deep attracor model.Meanwhile, we show it can be generalized well to more than one interfering speaker. Jun Wang 0091, Jie Chen 0057, Dan Su 0002, Lianwu Chen, Meng Yu 0003, Yanmin Qian, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2018 | Text-Dependent Speech Enhancement for Small-Footprint Robust Keyword Detection
Meng Yu 0003, Lianwu Chen, Jie Chen 0057, Jimeng Zheng, Dan Su 0002, Dong Yu 0001 |
INTERSPEECH | 1 |
| 2012 | A Triple-Microphone Real-Time Speech Enhancement Algorithm Based on Approximate Array Analytical Solutions
Meng Yu 0003, Ryan Ritch, Jack Xin |
INTERSPEECH | 1 |
| 2012 | Constrained Multichannel Speech Dereverberation
Meng Yu 0003, Frank K. Soong |
INTERSPEECH | 1 |
| 2012 | Exploring Off Time Nature for Speech Enhancement
Meng Yu 0003, Jack Xin |
INTERSPEECH | 1 |
| 2012 | Multi-Channel l1 Regularized Convex Speech Enhancement Model and Fast Computation by the Split Bregman MethodabstractA convex speech enhancement (CSE) method is presented based on convex optimization and pause detection of the speech sources. Channel spatial difference is identified for enhancing each speech source individually while suppressing other interfering sources. Sparse unmixing filters indicating channel spatial differences are sought byl1norm regularization and the split Bregman method. A subdivided split Bregman method is developed for efficiently solving the problem in severely reverberant environments. The speech pause detection is based on a binary mask source separation method. The CSE method is evaluated objectively and subjectively, and found to outperform a list of existing blind speech separation approaches on both synthetic and room recorded speech mixtures in terms of the overall computational speed and separation quality. Meng Yu 0003, Wenye Ma, Jack Xin, Stanley J. Osher |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Modeling Category Identification Using Sparse Instance Representation
Shunan Zhang, Michael D. Lee 0001, Meng Yu 0003, Jack Xin |
CogSci | 3 |
| 2010 | Reducing musical noise in blind source separation by time-domain sparse filters and split bregman methodabstractMusical noise often arises in the outputs of time-frequency binary mask based blind source separation approaches. Postprocessing is desired to enhance the separation quality. An efficient musical noise reduction method by time-domain sparse filters is presented using convex optimization. The sparse filters are sought by l1 regularization and the split Bregman method. The proposed musical noise reduction method is evaluated by both synthetic and room recorded speech and music data, and found to outperform existing musical noise reduction methods in terms of the objective and subjective measures. Index Terms: Musical noise, time-frequency mask, timedomain sparse filters, split Bregman method. Wenye Ma, Meng Yu 0003, Jack Xin, Stanley J. Osher |
INTERSPEECH | 2 |
| 2010 | Convexity and fast speech extraction by split bregman methodabstractA fast speech extraction (FSE) method is presented using convex optimization made possible by pause detection of the speech sources. Sparse unmixing filters are sought by l1 regularization and the split Bregman method. A subdivided split Bregman method is developed for efficiently estimating long reverberations in real room recordings. The speech pause detection is based on a binary mask source separation method. The FSE method is evaluated and found to outperform existing blind speech separation approaches on both synthetic and room recorded data in terms of the overall computational speed and separation quality. Index Terms: convexity, sparse filters, split Bregman method, fast blind speech extraction. Meng Yu 0003, Wenye Ma, Jack Xin, Stanley J. Osher |
INTERSPEECH | 1 |