VLDB 2026 Research / reviewers in the wild / expert
Yi Luo 0004
dblp:56/4437-4
· DBLP profile ↗
41ranked-venue papers
19as first author
25since 2021 · last 2025
0000-0002-7447-3885ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 14 first-author · 21 since 2021Artificial intelligence and machine learning · 19 · 10 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Apollo: Band-sequence Modeling for High-Quality Audio RestorationabstractAudio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of generative audio models necessitate high-fidelity audio. Typically, audio restoration is defined as a task of predicting undistorted audio from damaged input, often trained using a GAN framework to balance perception and distortion. Since audio degradation is primarily concentrated in mid- and high-frequency ranges, especially due to codecs, a key challenge lies in designing a generator capable of preserving low-frequency information while accurately reconstructing high-quality mid-and high-frequency content. Inspired by recent advancements in high-sample-rate music separation, speech enhancement, and audio codec models, we propose Apollo, a generative model designed for high-sample-rate audio restoration. Apollo employs an explicit frequency band split module to model the relationships between different frequency bands, allowing for more coherent and higher-quality restored audio. Evaluated on the MUSDB18-HQ and MoisesDB datasets, Apollo consistently outperforms existing SR-GAN models across various bit rates and music genres, particularly excelling in complex scenarios involving mixtures of multiple instruments and vocals. Apollo significantly improves music restoration quality while maintaining computational efficiency. The source code for Apollo is publicly available at https://github.com/JusperLee/Apollo. Yi Luo 0004 |
ICASSP | 2 |
| 2025 | WAKE: Watermarking Audio with Key Enrichment
Yaoxun Xu, Jianwei Yu 0001, Hangting Chen, Zhiyong Wu 0001, Xixin Wu, Dong Yu 0001, Rongzhi Gu, Yi Luo 0004 |
INTERSPEECH | 8 |
| 2024 | SECap: Speech Emotion Captioning with Large Language ModelabstractSpeech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human speech are often complex, and categorizing them into predefined groups can be insufficient to adequately represent speech emotions. On the contrary, describing speech emotions directly by means of natural language may be a more effective approach. Regrettably, there are not many studies available that have focused on this direction. Therefore, this paper proposes a speech emotion captioning framework named SECap, aiming at effectively describing speech emotions using natural language. Owing to the impressive capabilities of large language models in language comprehension and text generation, SECap employs LLaMA as the text decoder to allow the production of coherent speech emotion captions. In addition, SECap leverages HuBERT as the audio encoder to extract general speech features and Q-Former as the Bridge-Net to provide LLaMA with emotion-related speech features. To accomplish this, Q-Former utilizes mutual information learning to disentangle emotion-related speech features and speech contents, while implementing contrastive learning to extract more emotion-related speech features. The results of objective and subjective evaluations demonstrate that: 1) the SECap framework outperforms the HTSAT-BART baseline in all objective evaluations; 2) SECap can generate high-quality speech emotion captions that attain performance on par with human annotators in subjective mean opinion score tests. Yaoxun Xu, Hangting Chen, Jianwei Yu 0001, Qiaochu Huang, Zhiyong Wu 0001, Shixiong Zhang 0001, Guangzhi Li, Yi Luo 0004, Rongzhi Gu |
AAAI | 8 |
| 2024 | Improving Music Source Separation with Simo Stereo Band-Split RnnabstractWith the recent developments of novel neural network designs, the state-of-the-art of music source separation systems has been significantly advanced. For example, one of the recently-proposed models, the band-split RNN (BSRNN), proposed to split the input spectrogram into fine-grained subband features and perform interleaved sequence-level and band-level modeling, and achieved superior performance on the MUSDB18 dataset. However, the original BSRNN required different band-split schemes and networks for different instruments, which greatly increased the cost for model training and inference. Moreover, it was designed for single-channel scenario, while music signals are typically stereo. In this paper, we extend BSRNN to single-input-multi-output (SIMO) and stereo mode where all tracks are jointly extracted with a same network that supports stereo signal modeling. Experiment results show that the SIMO stereo BSRNN can effectively improve the overall performance of all tracks. Yi Luo 0004, Rongzhi Gu |
ICASSP | 1 |
| 2024 | Subnetwork-To-Go: Elastic Neural Network with Dynamic Training and Customizable InferenceabstractDeploying neural networks to different devices or platforms is in general challenging, especially when the model size is large or model complexity is high. Although there exist ways for model pruning or distillation, it is typically required to perform a full round of model training or finetuning procedure in order to obtain a smaller model that satisfies the model size or complexity constraints. Motivated by recent works on dynamic neural networks, we propose a simple way to train a large network and flexibly extract a subnetwork from it given a model size or complexity constraint during inference. We introduce a new way to allow a large model to be trained with dynamic depth and width during the training phase, and after the large model is trained we can select a subnetwork from it with arbitrary depth and width during the inference phase with a relatively better performance compared to training the subnetwork independently from scratch. Experiment results on a music source separation model show that our proposed method can effectively improve the separation performance across different subnetwork sizes and complexities with a single large model, and training the large model takes significantly shorter time than training all the different subnetworks. Yi Luo 0004 |
ICASSP | 2 |
| 2024 | AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech DataabstractRecently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that a considerable amount of the publicly available speech data is compromised by background noise, speech overlapping, lack of speech segmentation information, missing speaker labels, and incomplete transcriptions, which can largely hinder their usefulness. On the other hand, human annotation of speech data is both time-consuming and costly. To address this issue, we introduce an automatic in-the-wild speech data preprocessing framework (AutoPrep) in this paper, which is designed to enhance speech quality, generate speaker labels, and produce transcriptions automatically. The proposed AutoPrep framework comprises six components: speech enhancement, speech segmentation, speaker clustering, target speech extraction, quality filtering and automatic speech recognition. Experiments conducted on the open-sourced WenetSpeech and our self-collected AutoPrepWild corpora demonstrate that the proposed AutoPrep framework can generate preprocessed data with similar DNSMOS and PDNSMOS scores compared to several open-sourced TTS datasets. The corresponding TTS system can achieve up to 0.68 in-domain speaker similarity.1 Jianwei Yu 0001, Hangting Chen, Yanyao Bian, Yi Luo 0004, Jinchuan Tian, Mengyang Liu, Jiayi Jiang, Shuai Wang 0016 |
ICASSP | 5 |
| 2024 | ReZero: Region-Customizable Sound ExtractionabstractWe introduce region-customizable sound extraction (ReZero), a general and flexible framework for the multi-channel region-wise sound extraction (R-SE) task. R-SE task aims at extracting all active target sounds (e.g., human speech) within a specific, user-defined spatial region, which is different from conventional and existing tasks where a blind separation or a fixed, predefined spatial region are typically assumed. The spatial region can be defined as an angular window, a sphere, a cone, or other geometric patterns. Being a solution to the R-SE task, the proposed ReZero framework includes (1) definitions of different types of spatial regions, (2) methods for region feature extraction and aggregation, and (3) a multi-channel extension of the band-split RNN (BSRNN) model specified for the R-SE task. We design experiments for different microphone array geometries, different types of spatial regions, and comprehensive ablation studies on different system configurations. Experimental results on both simulated and real-recorded data demonstrate the effectiveness of ReZero. Demos are available athttps://innerselfm.github.io/rezero/. Rongzhi Gu, Yi Luo 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | On The Design and Training Strategies for Rnn-Based Online Neural Speech Separation SystemsabstractWhile the performance of offline neural speech separation systems has been greatly advanced by the recent development of novel neural network architectures, there is typically an inevitable performance gap between the systems and their online variants. In this paper, we investigate how RNN-based offline neural speech separation systems can be changed into their online counterparts while mitigating the performance degradation. We decompose or reorganize the forward and backward RNN layers in a bidirectional RNN layer to form an online path and an offline path, which enables the model to perform both online and offline processing with a same set of model parameters. We further introduce two training strategies for improving the online model via either a pretrained offline model or a multitask training objective. Experiment results show that compared to the online models that are trained from scratch, the proposed layer de-composition and reorganization schemes and training strategies can effectively mitigate the performance gap between two RNN-based offline separation models and their online variants. Yi Luo 0004 |
ICASSP | 2 |
| 2023 | TSpeech-AI System Description to the 5th Deep Noise Suppression (DNS) ChallengeabstractThis report presents the development of Tencent AI Lab’s personalized speech enhancement system for the 2023 ICASSP Signal Processing Grand Challenge – deep noise suppression (DNS) challenge1, which includes the use of a modified band-split recurrent neural network (BSRNN) and a multi-resolution spectrogram discriminator to improve perceptual quality metrics. The proposed system outperforms the baseline system by 0.047 and 0.036 final scores in the headset track and speaker phone track, respectively, and ranks among the top-32in both tracks. Jianwei Yu 0001, Hangting Chen, Yi Luo 0004, Rongzhi Gu, Chao Weng |
ICASSP | 3 |
| 2023 | Efficient Monaural Speech Enhancement with Universal Sample Rate Band-Split RNNabstractWhile recent developments on the design of neural networks have greatly advanced the state-of-the-art of speech enhancement and separation systems, practical applications of such networks often put extra constraints on their model size and computational complexity. Moreover, as different telecommunication services may have different transmission bandwidths which result in different signal sample rates, one model is typically designed for a particular sample rate. In this paper, we extend the usage of a recently proposed frequency-domain source separation model, the band-split RNN (BSRNN), to the task of universal-sample-rate resource efficient speech enhancement. BSRNN explicitly splits the spectrogram into different frequency bands and perform interleaved band-level and sequence-level modeling, and the bandwidths can be manually designed to balance the model size, computational cost, and performance. By properly designing the band-splitting scheme and the hyperparameters, a single BSRNN model can handle signals at a wide range of sample rates, and the computational cost required to process a lower-sample-rate signal can be smaller than that of a higher-sample-rate signal. Experiment results show that compared to various benchmark systems in speech enhancement and separation, our universal-sample-rate BSRNN (USR-BSRNN) achieves comparable or better signalto-noise ratio (SNR) performance at a same level of model size or computational cost. Jianwei Yu 0001, Yi Luo 0004 |
ICASSP | 2 |
| 2023 | FRA-RIR: Fast Random Approximation of the Image-source Method
Yi Luo 0004, Jianwei Yu 0001 |
INTERSPEECH | 1 |
| 2023 | Ultra Dual-Path Compression For Joint Echo Cancellation And Noise SuppressionabstractEcho cancellation and noise reduction are essential for full-duplex communication, yet most existing neural networks have high computational costs and are inflexible in tuning model complexity. In this paper, we introduce time-frequency dual-path compression to achieve a wide range of compression ratios on computational cost. Specifically, for frequency compression, trainable filters are used to replace manually designed filters for dimension reduction. For time compression, only using frame skipped prediction causes large performance degradation, which can be alleviated by a post-processing network with full sequence modeling. We have found that under fixed compression ratios, dual-path compression combining both the time and frequency methods will give further performance improvement, covering compression ratios from 4x to 32x with little model size change. Moreover, the proposed models show competitive performance compared with fast FullSubNet and DeepFilterNet. Hangting Chen, Jianwei Yu 0001, Yi Luo 0004, Rongzhi Gu, Zhuocheng Lu, Chao Weng |
INTERSPEECH | 3 |
| 2023 | High Fidelity Speech Enhancement with Band-split RNN
Jianwei Yu 0001, Hangting Chen, Yi Luo 0004, Rongzhi Gu, Chao Weng |
INTERSPEECH | 3 |
| 2023 | Music Source Separation With Band-Split RNNabstractThe performance of music source separation (MSS) models has been greatly improved in recent years thanks to the development of novel neural network architectures and training pipelines. However, recent model designs for MSS were mainly motivated by other audio processing tasks or other research fields, such as speech enhancement and speech separation, while the intrinsic characteristics and patterns of the music signals were not fully discovered. In this paper, we propose band-split RNN (BSRNN), a frequency-domain model that explictly splits the spectrogram of the mixture into subbands and perform interleaved band-level and sequence-level modeling. The choices of the bandwidths of the subbands can be determined by a priori knowledge or expert knowledge on the characteristics of the target source in order to optimize the performance on a certain type of target musical instrument. To better make use of unlabeled data, we also describe a semi-supervised model finetuning pipeline that can further improve the performance of the model. Experiment results show that BSRNN trained only on MUSDB18-HQ dataset significantly outperforms several top-ranking models in Music Demixing (MDX) Challenge 2021, and the semi-supervised finetuning stage further improves the performance on all four instrument tracks. Yi Luo 0004, Jianwei Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | A Time-Domain Real-Valued Generalized Wiener Filter for Multi-Channel Neural Separation SystemsabstractFrequency-domain beamformers have been successful in a wide range of multi-channel neural separation systems in the past years. However, the operations in conventional frequency-domain beamformers are typically independently-defined and complex-valued, which result in two drawbacks: the former does not fully utilize the advantage of end-to-end optimization, and the latter may introduce numerical instability during the training phase. Motivated by the recent success in end-to-end neural separation systems, in this paper we propose time-domain real-valued generalized Wiener filter (TD-GWF), a linear filter defined on a 2-D learnable real-valued signal transform. TD-GWF splits the transformed representation into groups and performs an minimum mean-square error (MMSE) estimation on all available channels on each of the groups. We show how TD-GWF can be connected to conventional filter-and-sum beamformers when certain signal transform and the number of groups are specified. Moreover, given the recent success in the sequential neural beamforming frameworks, we show how TD-GWF can be applied in such frameworks to perform iterative beamforming and separation to obtain an overall performance gain. Comprehensive experiment results show that TD-GWF performs consistently better than conventional frequency-domain beamformers in the sequential neural beamforming pipeline with various neural network architectures, microphone array scenarios, and task configurations. Yi Luo 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Dual-Path Modeling for Long Recording Speech Separation in MeetingsabstractThe continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording with a size-fixed window and process each window separately. Though effective, this extension fails to model the long dependency in speech and thus leads to sub-optimum performance. The recent proposed dual-path modeling could be a remedy to this problem, thanks to its capability in jointly modeling the cross-window dependency and the local-window processing. In this work, we further extend the dual-path modeling framework for CSS task. A transformer-based dual-path system is proposed, which integrates transform layers for global modeling. The proposed models are applied to LibriCSS, a real recorded multi-talk dataset, and consistent WER reduction can be observed in the ASR evaluation for separated speech. Also, a dual-path transformer equipped with convolutional layers is proposed. It significantly reduces the computation amount by 30% with better WER evaluation. Furthermore, the online processing dual-path models are investigated, which shows 10% relative WER reduction compared to the baseline. Chenda Li, Zhuo Chen 0006, Yi Luo 0004, Cong Han 0001, Tianyan Zhou, Keisuke Kinoshita, Marc Delcroix, Shinji Watanabe 0001, Yanmin Qian |
ICASSP | 3 |
| 2021 | Rethinking The Separation Layers In Speech Separation NetworksabstractModules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the numbers of input and output the same. While the majority of separation models only contain SIMO architectures, it has also been shown that certain two-stage separation systems integrated with a post-enhancement SISO module can improve the separation quality. Why performance improvements can be achieved by incorporating the SISO modules? Are SIMO modules always necessary? In this paper, we empirically examine those questions by designing models with varying configurations in the SIMO and SISO modules. We show that comparing with the standard SIMO-only design, a mixed SIMO-SISO design with a same model size is able to improve the separation performance especially under low-overlap conditions. We further validate the necessity of SIMO modules and show that SISO-only models are still able to perform separation without sacrificing the performance. The observations allow us to rethink the model design paradigm and present different views on how the separation is performed. Yi Luo 0004, Zhuo Chen 0006, Cong Han 0001, Chenda Li, Tianyan Zhou, Nima Mesgarani |
ICASSP | 1 |
| 2021 | Ultra-Lightweight Speech Separation Via Group CommunicationabstractModel size and complexity remain the biggest challenges in the deployment of speech enhancement and separation systems on low-resource devices such as earphones and hearing aids. Although methods such as compression, distillation and quantization can be applied to large models, they often come with a cost on the model performance. In this paper, we provide a simple model design paradigm that explicitly designs ultra-lightweight models without sacrificing the performance. Motivated by the sub-band frequency-LSTM (F-LSTM) architectures, we introduce the group communication (GroupComm), where a feature vector is split into smaller groups and a small processing block is used to perform inter-group communication. Unlike standard F-LSTM models where the sub-band outputs are concatenated, an ultra-small module is applied on all the groups in parallel, which allows a significant decrease on the model size. Experiment results show that comparing with a strong baseline model which is already lightweight, GroupComm can achieve on par performance with 35.6 times fewer parameters and 2.3 times fewer operations. Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
ICASSP | 1 |
| 2021 | Continuous Speech Separation Using Speaker Inventory for Long Recording
Cong Han 0001, Yi Luo 0004, Chenda Li, Tianyan Zhou, Keisuke Kinoshita, Shinji Watanabe 0001, Marc Delcroix, Hakan Erdogan, John R. Hershey, Nima Mesgarani, Zhuo Chen 0006 |
Interspeech | 2 |
| 2021 | Binaural Speech Separation of Moving Speakers With Preserved Spatial Cues
Cong Han 0001, Yi Luo 0004, Nima Mesgarani |
Interspeech | 2 |
| 2021 | Empirical Analysis of Generalized Iterative Speech Separation Networks
Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
Interspeech | 1 |
| 2021 | Distortion-Controlled Training for end-to-end Reverberant Speech Separation with Auxiliary Autoencoding LossabstractThe performance of speech enhancement and separation systems in anechoic environments has been significantly advanced with the recent progress in end-to-end neural network architectures. However, the performance of such systems in reverberant environments is yet to be explored. A core problem in reverberant speech separation is about the training and evaluation metrics. Standard time-domain metrics may introduce unexpected distortions during training and fail to properly evaluate the separation performance due to the presence of the reverberations. In this paper, we first introduce the "equal-valued contour" problem in reverberant separation where multiple outputs can lead to the same performance measured by the common metrics. We then investigate how "better" outputs with lower target-specific distortions can be selected by auxiliary autoencoding training (A2T). A2T assumes that the separation is done by a linear operation on the mixture signal, and it adds an loss term on the autoencoding of the direct-path target signals to ensure that the distortion introduced on the direct-path signals is controlled during separation. Evaluations on separation signal quality and speech recognition accuracy show that A2T is able to control the distortion on the direct-path signals and improve the recognition accuracy. Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
SLT | 1 |
| 2021 | Dual-Path RNN for Long Recording Speech SeparationabstractContinuous speech separation (CSS) is an arising task in speech separation aiming at separating overlap-free targets from a long, partially-overlapped recording. A straightforward extension of previously proposed sentence-level separation models to this task is to segment the long recording into fixed-length blocks and perform separation on them independently. However, such simple extension does not fully address the cross-block dependencies and the separation performance may not be satisfactory. In this paper, we focus on how the block-level separation performance can be improved by exploring methods to utilize the cross-block information. Based on the recently proposed dual-path RNN (DPRNN) architecture, we investigate how DPRNN can help the block-level separation by the interleaved intra- and inter-block modules. Experiment results show that DPRNN is able to significantly outperform the baseline block-level model in both offline and block-online configurations under certain settings. Chenda Li, Yi Luo 0004, Cong Han 0001, Jinyu Li 0001, Takuya Yoshioka, Tianyan Zhou, Marc Delcroix, Keisuke Kinoshita, Christoph Böddeker, Yanmin Qian, Shinji Watanabe 0001, Zhuo Chen 0006 |
SLT | 2 |
| 2021 | Integration of Speech Separation, Diarization, and Recognition for Multi-Speaker Meetings: System Description, Comparison, and AnalysisabstractMulti-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, speaker diarization, and automatic speech recognition (ASR) in the last decade, it has become possible to build pipelines that achieve reasonable error rates on this task. In this paper, we propose an end-to-end modular system for the LibriCSS meeting data, which combines independently trained separation, diarization, and recognition components, in that order. We study the effect of different state-of-the-art methods at each stage of the pipeline, and report results using task-specific metrics like SDR and DER, as well as downstream WER. Experiments indicate that the problem of overlapping speech for diarization and ASR can be effectively mitigated with the presence of a well-trained separation module. Our best system achieves a speaker-attributed WER of 12.7%, which is close to that of a non-overlapping ASR. Desh Raj, Pavel Denisov, Zhuo Chen 0006, Hakan Erdogan, Zili Huang, Maokui He, Shinji Watanabe 0001, Jun Du 0002, Takuya Yoshioka, Yi Luo 0004, Naoyuki Kanda, Jinyu Li 0001, Scott Wisdom, John R. Hershey |
SLT | 10 |
| 2021 | Group Communication With Context Codec for Lightweight Source SeparationabstractDespite the recent progress on neural network architectures for speech separation, the balance between the model size, model complexity and model performance is still an important and challenging problem for the deployment of such models to low-resource platforms. In this paper, we propose two simple modules, group communication and context codec, that can be easily applied to a wide range of architectures to jointly decrease the model size and complexity without sacrificing the performance. A group communication module splits a high-dimensional feature into groups of low-dimensional features and captures the inter-group dependency. A separation module with a significantly smaller model size can then be shared by all the groups. A context codec module, containing a context encoder and a context decoder, is designed as a learnable downsampling and upsampling module to decrease the length of a sequential feature processed by the separation module. The combination of the group communication and the context codec modules is referred to as the GC3 design. Experimental results show that applying GC3 on multiple network architectures for speech separation can achieve on-par or better performance with as small as 2.5% model size and 17.6% model complexity, respectively. Yi Luo 0004, Cong Han 0001, Nima Mesgarani |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Continuous Speech Separation: Dataset and AnalysisabstractThis paper describes a dataset and protocols for evaluating continuous speech separation algorithms. Most prior speech separation studies use pre-segmented audio signals, which are typically generated by mixing speech utterances on computers so that they fully overlap. Also, the separation algorithms have often been evaluated based on signal-based metrics such as signal-to-distortion ratio. However, in natural conversations, speech signals are continuous and contain both overlapped and overlap-free regions. In addition, the signal-based metrics only have weak correlation with automatic speech recognition (ASR) accuracy. Not only does this make it hard to assess the practical relevance of the tested algorithms, it also hinders researchers from developing systems that can be readily applied to real scenarios. In this paper, we define continuous speech separation (CSS) as a task of generating a set of non-overlapped speech signals from a continuous audio stream that contains multiple utterances that are partially overlapped by a varying degree. A new real recording dataset, called LibriCSS, is derived from LibriSpeech by concatenating the corpus utterances to simulate conversations and capturing the audio replays with far-field microphones. A Kaldi-based ASR evaluation protocol is established by using a well-trained multi-conditional acoustic model. A recently proposed speaker-independent CSS algorithm is investigated by using LibriCSS. The dataset and evaluation scripts are made available to facilitate the research in this direction1. Zhuo Chen 0006, Takuya Yoshioka, Liang Lu 0001, Tianyan Zhou, Zhong Meng, Yi Luo 0004, Jian Wu 0027, Jinyu Li 0001 |
ICASSP | 6 |
| 2020 | Real-Time Binaural Speech Separation with Preserved Spatial CuesabstractDeep learning speech separation algorithms have achieved great success in improving the quality and intelligibility of separated speech from mixed audio. Most previous methods focused on generating a single-channel output for each of the target speakers, hence discarding the spatial cues needed for the localization of sound sources in space. However, preserving the spatial information is important in many applications that aim to accurately render the acoustic scene such as in hearing aids and augmented reality (AR). Here, we propose a speech separation algorithm that preserves the interaural cues of separated sound sources and can be implemented with low latency and high fidelity, therefore enabling a real-time modification of the acoustic scene. Based on the time-domain audio separation network (TasNet), a single-channel time-domain speech separation system that can be implemented in real-time, we propose a multi-input-multi-output (MIMO) end-to-end extension of TasNet that takes binaural mixed audio as input and simultaneously separates target speakers in both channels. Experimental results show that the proposed end-to-end MIMO system is able to significantly improve the separation performance and keep the perceived location of the modified sources intact in various acoustic scenes. Cong Han 0001, Yi Luo 0004, Nima Mesgarani |
ICASSP | 2 |
| 2020 | End-to-end Microphone Permutation and Number Invariant Multi-channel Speech SeparationabstractAn important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones. The former requires the system to be invariant to different indexing of the microphones with the same locations, while the latter requires the system to be able to process inputs with varying dimensions. Conventional optimization-based beamforming techniques satisfy these requirements by definition, while for deep learning-based end-to-end systems those constraints are not fully addressed. In this paper, we propose transform-average-concatenate (TAC), a simple design paradigm for channel permutation and number invariant multi-channel speech separation. Based on the filter-and-sum network (FaSNet), a recently proposed end-to-end time-domain beamforming system, we show how TAC significantly improves the separation performance across various numbers of microphones in noisy reverberant separation tasks with ad-hoc arrays. Moreover, we show that TAC also significantly improves the separation performance with fixed geometry array configuration, further proving the effectiveness of the proposed paradigm in the general problem of multi-microphone speech separation. Yi Luo 0004, Zhuo Chen 0006, Nima Mesgarani, Takuya Yoshioka |
ICASSP | 1 |
| 2020 | Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech SeparationabstractRecent studies in deep learning-based speech separation have proven the superiority of time-domain approaches to conventional time-frequency-based methods. Unlike the time-frequency domain approaches, the time-domain separation systems often receive input sequences consisting of a huge number of time steps, which introduces challenges for modeling extremely long sequences. Conventional recurrent neural networks (RNNs) are not effective for modeling such long sequences due to optimization difficulties, while one-dimensional convolutional neural networks (1-D CNNs) cannot perform utterance-level sequence modeling when its receptive field is smaller than the sequence length. In this paper, we propose dual-path recurrent neural network (DPRNN), a simple yet effective method for organizing RNN layers in a deep structure to model extremely long sequences. DPRNN splits the long sequential input into smaller chunks and applies intra- and inter-chunk operations iteratively, where the input length can be made proportional to the square root of the original sequence length in each operation. Experiments show that by replacing 1-D CNN with DPRNN and apply sample-level modeling in the time-domain audio separation network (TasNet), a new state-of-the-art performance on WSJ0-2mix is achieved with a 20 times smaller model than the previous best system. Yi Luo 0004, Zhuo Chen 0006, Takuya Yoshioka |
ICASSP | 1 |
| 2020 | Separating Varying Numbers of Sources with Auxiliary Autoencoding LossabstractMany recent source separation systems are designed to separate a fixed number of sources out of a mixture. In the cases where the source activation patterns are unknown, such systems have to either adjust the number of outputs or to identify invalid outputs from the valid ones. Iterative separation methods have gain much attention in the community as they can flexibly decide the number of outputs, however (1) they typically rely on long-term information to determine the stopping time for the iterations, which makes them hard to operate in a causal setting; (2) they lack a fault tolerance mechanism when the estimated number of sources is different from the actual number. In this paper, we propose a simple training method, the auxiliary autoencoding permutation invariant training (A2PIT), to alleviate the two issues. A2PIT assumes a fixed number of outputs and uses auxiliary autoencoding loss to force the invalid outputs to be the copies of the input mixture, and detects invalid outputs in a fully unsupervised way during inference phase. Experiment results show that A2PIT is able to improve the separation performance across various numbers of speakers and effectively detect the number of speakers in a mixture. Yi Luo 0004, Nima Mesgarani |
INTERSPEECH | 1 |
| 2020 | An End-to-End Architecture of Online Multi-Channel Speech SeparationabstractAlthough mask based adaptive beamforming technique benefits speech recognition in far-field, noisy and multi-talker scenarios, it depends on the long time context to estimate target and interference statistics, thus when applied in applications with low latency requirement, its performance usually drops drastically. In contrast, the fixed beamformers do not import time delay but usually have limited capability in acoustic cancellation of interfering source. In this work, we propose a novel multi-channel speech separation system that targets at overlapped speech recognition with low latency processing, which includes four jointly optimized components: a pre-separator, a set of fixed beamformer, an attentional selection module and neural post filtering. With proposed model, low latency processing is achieved by utilizing the known microphone geometry information, while keeps the high quality separation through neural post filtering and end-to-end optimization. In our experiments, we show that the proposed system achieves comparable performance in offline evaluation with the mask based MVDR and speech extraction system, while yield remarkable improvements in the online evaluation. Jian Wu 0027, Zhuo Chen 0006, Jinyu Li 0001, Takuya Yoshioka, Zhili Tan, Ed Lin, Yi Luo 0004, Lei Xie 0001 |
INTERSPEECH | 7 |
| 2019 | FaSNet: Low-Latency Adaptive Beamforming for Multi-Microphone Audio ProcessingabstractBeamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called neural beamformers, have achieved significant improvements in both signal quality (e.g. signal-to-noise ratio (SNR)) and speech recognition (e.g. word error rate (WER)). Such systems are generally non-causal and require a large context for robust estimation of inter-channel features, which is impractical in applications requiring low-latency responses. In this paper, we propose filter-and-sum network (FaSNet), a time-domain, filter-based beamforming approach suitable for low-latency scenarios. FaSNet has a two-stage system design that first learns frame-level time-domain adaptive beamforming filters for a selected reference channel, and then calculate the filters for all remaining channels. The filtered outputs at all channels are summed to generate the final output. Experiments show that despite its small model size, FaSNet is able to outperform several traditional oracle beamformers with respect to scale-invariant signal-to-noise ratio (SI-SNR) in reverberant speech enhancement and separation tasks. Moreover, when trained with a frequency-domain objective function on the CHiME-3 dataset, FaSNet achieves 14.3% relative word error rate reduction (RWERR) compared with the baseline model. These results show the efficacy of FaSNet particularly in reverberant and noisy signal conditions. Yi Luo 0004, Cong Han 0001, Nima Mesgarani, Enea Ceolini, Shih-Chii Liu |
ASRU | 1 |
| 2019 | Augmented Time-frequency Mask Estimation in Cluster-based Source Separation AlgorithmsabstractTime-frequency mask estimation with various clustering approaches has proven effective in solving the audio source separation problem. In this framework, the time-frequency bins of the mixture spectrogram are represented in a high-dimensional embedding space, where various methods can be applied to group the embedded points to calculate either hard or soft source assignments and subsequently the time-frequency masks. However, the mismatch between the assignment algorithm during the training and inference phases in majority of the current approaches leads to a suboptimal solution, because the assignment objective that is used during the training (e.g. ideal binary mask) is not the same as the one used during the inference phase (e.g. k-means clustering). We propose a method to reduce the mismatch between these two conditions where the source embedding is trained such that the source assignment during training and inference phases results in similar outcomes. Our results show that matching the source assignment during training- and inference-phase results in more accurate and consistent mask estimation in the inference phase which significantly improves the source separation accuracy for various hard and soft clustering methods. Yi Luo 0004, Nima Mesgarani |
ICASSP | 1 |
| 2019 | Online Deep Attractor Network for Real-time Single-channel Speech SeparationabstractSpeaker-independent speech separation is a challenging audio processing problem. In recent years, several deep learning algorithms have been proposed to address this problem. The majority of these methods use noncausal implementation which limits their application in real-time scenarios such as in wearable hearing devices and low-latency telecommunication. In this paper, we propose the Online Deep Attractor Network (ODANet), an extension to the Deep Attractor Network (DANet) which is causal and enables real-time speech separation. In contrast with DANet that estimates the global attractor point for each speaker using the entire utterance, ODANet estimates the attractors for each time step and tracks them using a dynamic weighting function with only causal information. This not only solves the speaker tracking problem, but also allows ODANet to generate more stable embeddings across time. Experimental results show that ODANet can achieve a similar separation accuracy as the noncausal DANet in both two speaker and three speaker speech separation problems, which makes it a suitable candidate for applications that require robust real-time speech processing. Cong Han 0001, Yi Luo 0004, Nima Mesgarani |
ICASSP | 2 |
| 2019 | Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech SeparationabstractSingle-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous methods have formulated the separation problem through the time-frequency representation of the mixed signal, which has several drawbacks, including the decoupling of the phase and magnitude of the signal, the suboptimality of time-frequency representation for speech separation, and the long latency of the entire system. To address these shortcomings, we propose a fully-convolutional time-domain audio separation network (Conv-TasNet), a deep learning framework for end-to-end time-domain speech separation. Conv-TasNet uses a linear encoder to generate a representation of the speech waveform optimized for separating individual speakers. Speaker separation is achieved by applying a set of weighting functions (masks) to the encoder output. The modified encoder representations are then inverted back to the waveforms using a linear decoder. The masks are found using a temporal convolutional network (TCN) consisting of stacked 1-D dilated convolutional blocks, which allows the network to model the long-term dependencies of the speech signal while maintaining a small model size. The proposed Conv-TasNet system significantly outperforms previous time-frequency masking methods in separating two- and three-speaker mixtures. Additionally, Conv-TasNet surpasses several ideal time-frequency magnitude masks in two-speaker speech separation as evaluated by both objective distortion measures and subjective quality assessment by human listeners. Finally, Conv-TasNet has a significantly smaller model size and a much shorter minimum latency, making it a suitable solution for both offline and real-time speech separation applications. This study therefore represents a major step toward the realization of speech separation systems for real-world speech processing technologies. Yi Luo 0004, Nima Mesgarani |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | TaSNet: Time-Domain Audio Separation Network for Real-Time, Single-Channel Speech SeparationabstractRobust speech processing in multi-talker environments requires effective speech separation. Recent deep learning systems have made significant progress toward solving this problem, yet it remains challenging particularly in real-time, short latency applications. Most methods attempt to construct a mask for each source in time-frequency representation of the mixture signal which is not necessarily an optimal representation for speech separation. In addition, time-frequency decomposition results in inherent problems such as phase/magnitude decoupling and long time window which is required to achieve sufficient frequency resolution. We propose Time-domain Audio Separation Network (TasNet) to overcome these limitations. We directly model the signal in the time-domain using an encoder-decoder framework and perform the source separation on nonnegative encoder outputs. This method removes the frequency decomposition step and reduces the separation problem to estimation of source masks on encoder outputs which is then synthesized by the decoder. Our system outperforms the current state-of-the-art causal and noncausal speech separation algorithms, reduces the computational cost of speech separation, and significantly reduces the minimum required latency of the output. This makes TasNet suitable for applications where low-power, real-time implementation is desirable such as in hearable and telecommunication devices. Yi Luo 0004, Nima Mesgarani |
ICASSP | 1 |
| 2018 | Real-time Single-channel Dereverberation and Separation with Time-domain Audio Separation Network
Yi Luo 0004, Nima Mesgarani |
INTERSPEECH | 1 |
| 2018 | Music Source Activity Detection and Separation Using Deep Attractor Network
Rajath Kumar, Yi Luo 0004, Nima Mesgarani |
INTERSPEECH | 2 |
| 2018 | Speaker-Independent Speech Separation With Deep Attractor NetworkabstractDespite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and masker speakers in the mixture (permutation problem), and the second is the unknown number of speakers in the mixture (output dimension problem). We propose a novel deep learning framework for speech separation that addresses both of these issues. We use a neural network to project the time-frequency representation of the mixture signal into a high-dimensional embedding space. A reference point (attractor) is created in the embedding space to represent each speaker which is defined as the centroid of the speaker in the embedding space. The time-frequency embeddings of each speaker are then forced to cluster around the corresponding attractor point which is used to determine the time-frequency assignment of the speaker. We propose three methods for finding the attractors for each source in the embedding space and compare their advantages and limitations. The objective function for the network is standard signal reconstruction error which enables end-to-end operation during both training and test phases. We evaluated our system using the Wall Street Journal dataset (WSJ0) on two and three speaker mixtures and report comparable or better performance than other state-of-the-art deep learning methods for speech separation. Yi Luo 0004, Zhuo Chen 0006, Nima Mesgarani |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Deep attractor network for single-microphone speaker separationabstractDespite the overwhelming success of deep learning in various speech processing tasks, the problem of separating simultaneous speakers in a mixture remains challenging. Two major difficulties in such systems are the arbitrary source permutation and unknown number of sources in the mixture. We propose a novel deep learning framework for single channel speech separation by creating attractor points in high dimensional embedding space of the acoustic signals which pull together the time-frequency bins corresponding to each source. Attractor points in this study are created by finding the centroids of the sources in the embedding space, which are subsequently used to determine the similarity of each bin in the mixture to each source. The network is then trained to minimize the reconstruction error of each source by optimizing the embeddings. The proposed model is different from prior works in that it implements an end-to-end training, and it does not depend on the number of sources in the mixture. Two strategies are explored in the test time, K-means and fixed attractor points, where the latter requires no post-processing and can be implemented in real-time. We evaluated our system on Wall Street Journal dataset and show 5.49% improvement over the previous state-of-the-art methods. Zhuo Chen 0006, Yi Luo 0004, Nima Mesgarani |
ICASSP | 2 |
| 2017 | Deep clustering and conventional networks for music separation: Stronger togetherabstractDeep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging situations such as music source separation. Contrary to conventional networks that directly estimate the source signals, deep clustering generates an embedding for each time-frequency bin, and separates sources by clustering the bins in the embedding space. We show that deep clustering outperforms conventional networks on a singing voice separation task, in both matched and mismatched conditions, even though conventional networks have the advantage of end-to-end training for best signal approximation, presumably because its more flexible objective engenders better regularization. Since the strengths of deep clustering and conventional network architectures appear complementary, we explore combining them in a single hybrid network trained via an approach akin to multi-task learning. Remarkably, the combination significantly outperforms either of its components. Yi Luo 0004, Zhuo Chen 0006, John R. Hershey, Jonathan Le Roux, Nima Mesgarani |
ICASSP | 1 |